Scientists Expose AI Reasoning Leak Across 3 Major Providers, Raising Distillation Fears
Updated
Updated · WIRED · Aug 11
Scientists Expose AI Reasoning Leak Across 3 Major Providers, Raising Distillation Fears
3 articles · Updated · WIRED · Aug 11
Summary
Researchers showed that hidden reasoning traces from OpenAI, Anthropic and Google models can be extracted by replaying encrypted outputs through smaller related models with weaker safeguards.
That method also recovered sensitive data such as API keys and passwords from captured traces, prompting all three companies to change their APIs; the direct personal-data leak is now mitigated.
Tests on 90 questions found Moonshot AI’s open-weight Kimi K3 sometimes produced strikingly similar continuations to hidden traces from Claude Opus 4.8 and GPT 5.6 Sol, though the paper says this cannot prove distillation.
The work suggests closed-model APIs may reveal more training value than companies assumed, and fully stopping reasoning extraction would require deeper architectural changes rather than a simple patch.
Distillation has become a US-China flashpoint as lawmakers and companies debate whether restricting it would curb Chinese gains or slow AI progress more broadly.
Could the discovery of leaked reasoning traces finally prove that popular open-source AI models are secretly copying proprietary tech?
If smaller AI models can expose the hidden thoughts of advanced systems, are your company's deepest secrets already compromised?
When security through obscurity fails in AI, how can enterprises protect their data from being weaponized by the models themselves?
The 2026 AI Distillation Crisis: 16 Million Exchanges, U.S.-China Espionage, and the Battle for Global AI Control
Overview
In 2026, the global AI race erupted when Anthropic accused major Chinese labs of using massive proxy networks to distill capabilities from its Claude models, bypassing U.S. export controls and hardware restrictions. This sparked a wave of U.S. government action, including threats of sanctions and new export controls targeting AI models themselves. Forensic evidence, like internal build strings found in Chinese models, fueled the controversy, but proving theft remains complex. The dispute exposed deep industry divides over open-source AI and accusations of hypocrisy, as U.S. labs had previously scraped data under 'fair use.' Ultimately, the controversy accelerated U.S.-China tech decoupling and reshaped global AI governance debates.