OpenAI’s custom silicon is officially taking on Nvidia. The newly detailed OpenAI Jalapeño chip has posted benchmark results that outperform Nvidia’s GB200 and GB300 superchips in AI inference tasks, delivering faster responses and significantly higher energy efficiency. This development is critical for enterprise developers and researchers who rely on high-throughput models to power real-time applications.
By drastically reducing the time between tokens (TBT), Jalapeño aims to make AI agents more responsive while lowering the massive energy costs associated with running large language models at scale. The custom hardware signals a major shift in how the AI giant plans to handle the daily operational load of its most popular services.
Breaking the Latency and Throughput Trade-off
First introduced in June, Jalapeño is an Application-Specific Integrated Circuit (ASIC) developed in partnership with Broadcom. It is purpose-built exclusively for AI inference - the process of running a trained model to generate responses or deploy an agent. According to Richard Ho, OpenAI's hardware vice president, the chip offers the "best of both worlds" by eliminating the traditional compromise where systems "have to make a trade-off" between low latency and high throughput.
To prove these claims, the company published official benchmark results using the InferenceX platform. The tests pitted Jalapeño against the best recorded results from Nvidia's top-tier hardware across heavy-duty models like GPT-OSS 120B, DeepSeek R1, and Kimi K2.5 1T. The custom ASIC delivered 1.5 to 1.9 times more AI work per watt compared to the Nvidia systems.
Furthermore, Jalapeño achieved 1.7 to 3.6 times lower end-to-end latency across all three tested models. Ho explained that this performance leap translates directly to "faster responses, more responsive agents, and more reliable access as the demand grows" for end users.
Deployment Timeline and Nvidia Partnership
OpenAI plans to deploy Jalapeño in "small volumes" by the end of 2026, with a broader strategy to "ramp the volume up" throughout 2027. However, the company has not disclosed the exact number of chips it intends to integrate into its data centers next year. Development on the second and third generations of the silicon is already underway.
Despite the impressive benchmarks, OpenAI is not abandoning its current infrastructure. Ho confirmed that the company does not expect to replace its entire hardware lineup with Jalapeño. Instead, OpenAI's overarching compute strategy will continue to rely on "very good partners" like Nvidia, indicating a hybrid approach to its massive processing needs.
The Economics of Inference Are Shifting
OpenAI’s transition from a pure software company to a custom silicon designer is a necessary evolution to control its own operational destiny. While Nvidia completely dominates the training phase of AI models, the real long-term financial burden lies in inference - serving millions of user queries and autonomous agent actions every single day. By achieving up to 1.9x better energy efficiency, Jalapeño directly attacks the primary bottleneck of scaling AI businesses: the sheer cost of electricity and compute time.
However, keeping Nvidia as a core partner is a strategic necessity rather than just corporate diplomacy. OpenAI still requires Nvidia's massive GPU clusters to train the next generation of foundational models. By offloading the daily, repetitive inference workloads to its own highly efficient ASICs, OpenAI can reserve its expensive Nvidia hardware strictly for heavy-duty training, optimizing its capital expenditure while improving the end-user experience.