Researchers Find 88.7% of AI Queries Run Locally, Cutting Cost and Energy by 60-80%
Updated
Updated · michaeljburry.substack.com · Aug 21
Researchers Find 88.7% of AI Queries Run Locally, Cutting Cost and Energy by 60-80%
1 articles · Updated · michaeljburry.substack.com · Aug 21
Summary
A revised August 7 paper measuring “intelligence per watt” across 1 million-plus queries, 20-plus models and eight accelerators found 88.7% of single-turn chat and reasoning tasks can be handled by small local models.
The study said local AI efficiency improved 5.3 times from 2023 to 2025, driven by a 3.1 times model gain and a 1.7 times hardware gain.
Hybrid routing between local devices and cloud models reduced energy, compute and cost by 60% to 80% versus a batched cloud baseline, while an 80%-accurate router captured about 80% of ideal gains without hurting answer quality.
Those results challenge assumptions behind nonstop data-center expansion by suggesting many enterprise AI workloads may shift toward cheaper small models running locally, with cloud systems reserved for harder queries.