SemiAnalysis Releases AgentX 1.0 Benchmark at 1 Million Context After $3 Million Build
Updated
Updated · newsletter.semianalysis.com · Aug 24
SemiAnalysis Releases AgentX 1.0 Benchmark at 1 Million Context After $3 Million Build
1 articles · Updated · newsletter.semianalysis.com · Aug 24
Summary
AgentX 1.0 is now open-sourced under Apache 2.0 as a multi-turn agentic coding benchmark built for 1 million-token context, with SemiAnalysis saying fixed-sequence tests no longer reflect production inference.
More than $3 million and about 2 MW of continuous compute across 1,000-plus chips went into the dataset and test matrix, which replays realistic long-context sessions with subagents, tool calls, KV-cache reuse and offload.
The benchmark arrives as agentic workloads surge: SemiAnalysis says they now dominate production inference traffic, and OpenAI enterprise agent spending overtook ChatGPT spending in April 2026.
AgentX has already driven 70-plus upstream optimization pull requests across vLLM, SGLang, TensorRT-LLM, ATOM, Dynamo, LMCache and Mooncake, with SemiAnalysis arguing many gains transfer directly to live deployments.
Initial results show both Nvidia and AMD performing strongly on some frontier models, while SemiAnalysis plans an update in three to four weeks with more optimizations, added hardware including TPUs, and refreshed comparisons.
If agentic AI is a systems problem, could AgentX's new metrics render today's fastest AI chips obsolete by exposing critical memory bottlenecks?
With 50 optimization fixes already driven by AgentX, what hidden flaws in our current AI infrastructure is this benchmark quietly exposing?
Since AgentX relies on deterministic replays, could optimizing for it accidentally create a massive blind spot for unpredictable real-world AI failures?
Benchmarking the Future: AgentX 1.0’s $3 Million, Million-Token Standard and Its Impact on AI Infrastructure and Sustainability
Overview
Following the November 2025 Claude Code inflection point, long-context, multi-turn agentic workloads rapidly grew and soon overtook traditional AI usage. This shift exposed the limits of old evaluation tools, prompting SemiAnalysis to invest over $3 million to create AgentX, the first open-source, multi-turn agentic coding benchmark. Released in August 2026, AgentX replaced outdated single-turn benchmarks and was quickly adopted in InferenceXv3, leading to major improvements in how AI hardware and software are measured. The benchmark’s release sparked a wave of software optimizations, especially in memory management, and intensified the competition between NVIDIA and AMD as the industry raced to meet the demands of modern agentic AI.