Clean Scores, Buried Evidence, and Confident Wrong: A Receipt-Based Audit of Frontier Agentic QA
Updated
Updated · arxiv.org · Sep 14
Clean Scores, Buried Evidence, and Confident Wrong: A Receipt-Based Audit of Frontier Agentic QA
1 articles · Updated · arxiv.org · Sep 14
Summary
A new audit reveals that leading AI models struggle with accuracy when evidence is buried in complex, document-heavy tasks.
The study found that models produced more errors, forced declarations, and higher costs when evidence was less accessible, despite high confidence scores.
Researchers warn that claim-level evidence receipts and rigorous human verification are necessary, especially for regulated sectors like finance and defense.