OpenAI Says Jalapeño Chip Delivers 1.5-1.9x More AI Work per Watt
Updated
Updated · OpenAI · Aug 25
OpenAI Says Jalapeño Chip Delivers 1.5-1.9x More AI Work per Watt
3 articles · Updated · OpenAI · Aug 25
Summary
1.5-1.9x more AI work per watt and 1.7-3.6x lower end-to-end latency were recorded for OpenAI’s Jalapeño across GPT-OSS 120B, DeepSeek R1 670B and Kimi K2.5 1T.
InferenceX benchmark tests measured matched user experience across throughput and low-latency settings, with OpenAI saying Jalapeño stayed on the Pareto frontier by improving speed, efficiency and responsiveness at once.
550 watts or less of sustained power was measured on tested workloads, versus Jalapeño’s 700-watt rating, while highly interactive workloads showed 2.1-4.1x higher performance.
Nine months from design to tapeout, the chip was developed with AI assistance, and OpenAI said AI-generated code sped selected model blocks by 1.5-1.8x over human-written versions.
By year-end, OpenAI plans to start deploying Jalapeño in its own infrastructure as the first step in a multigenerational inference-chip roadmap alongside continued use of Nvidia and other partners’ accelerators.
With Jalapeño cutting inference costs by 50%, will OpenAI's captive AI chip permanently disrupt Nvidia's dominance in the hardware market?
Since Jalapeño is strictly captive silicon, how will independent developers ever match the lightning-fast inference speeds of OpenAI's closed ecosystem?
If AI models helped design OpenAI's new chip in just nine months, could future silicon completely eliminate human hardware engineers?