Updated
Updated · michaeljburry.substack.com · Aug 21
Researchers Find 88.7% of AI Queries Run Locally, Cutting Cost and Energy by 60-80%
Updated
Updated · michaeljburry.substack.com · Aug 21

Researchers Find 88.7% of AI Queries Run Locally, Cutting Cost and Energy by 60-80%

1 articles · Updated · michaeljburry.substack.com · Aug 21

Summary

  • A revised August 7 paper measuring “intelligence per watt” across 1 million-plus queries, 20-plus models and eight accelerators found 88.7% of single-turn chat and reasoning tasks can be handled by small local models.
  • The study said local AI efficiency improved 5.3 times from 2023 to 2025, driven by a 3.1 times model gain and a 1.7 times hardware gain.
  • Hybrid routing between local devices and cloud models reduced energy, compute and cost by 60% to 80% versus a batched cloud baseline, while an 80%-accurate router captured about 80% of ideal gains without hurting answer quality.
  • Those results challenge assumptions behind nonstop data-center expansion by suggesting many enterprise AI workloads may shift toward cheaper small models running locally, with cloud systems reserved for harder queries.

Insights

Could the sudden rise of hyper-efficient local AI models trigger a massive crash in centralized data center investments?
If AI infrastructure is single-handedly propping up the U.S. economy, what happens when the power grid finally maxes out?
Are tech giants masking a broader economic recession by using massive cash reserves to build rate-proof AI empires?