OpenAI unveiled the first official benchmarks for Jalapeño, its custom inference chip developed with Broadcom, at the Hot Chips conference on August 25, 2026. The company reported that on SemiAnalysis’s public InferenceX benchmark, Jalapeño delivered between 1.5 and 1.9 times more work per watt than Nvidia’s Blackwell GB200 and GB300 rack systems, while cutting end-to-end response times by 1.7 to 3.6 times.
The tests covered three open models: OpenAI’s GPT-OSS 120B, DeepSeek’s R1 670B, and Moonshot AI’s Kimi K2.5 1T. For highly interactive workloads that require frequent back-and-forth, OpenAI said the advantage widened to 2.1 to 4.1 times faster performance. On Kimi K2.5, the largest model tested, Jalapeño reached roughly 1.5 times the peak performance per watt and 3.4 times lower latency.
Richard Ho, OpenAI’s head of hardware, said the chip handles more work at once while also responding faster—something many chips cannot do simultaneously. Although the chip is rated at 700 watts, OpenAI said it stayed at or below 550 watts on the tested workloads.
OpenAI does not plan to sell or rent Jalapeño computing capacity. Ho said the company is “struggling to have enough” compute for its own needs and is limited by data center power rather than budget or floor space. The chip will be deployed in racks of 128, with a full pod containing 2,048 ASICs, delivering 1.7 exaflops of 4-bit compute and 27.5 terabytes of HBM4.
OpenAI said it used its own models to accelerate the design process, moving from design to tape-out in just nine months. A second-generation chip is already “deep into development,” and a third is underway. Jalapeño will be deployed in small volumes by the end of 2026 before broader scaling through 2027.
Despite the benchmark claims, OpenAI will not abandon its GPU fleet. Ho described Jalapeño as one piece of a compute strategy that still relies on “very, very good partners” at Nvidia and Cerebras. SemiAnalysis verified the InferenceX runs in person but noted that all figures were supplied by OpenAI, and the chip has not undergone independent testing. Sam Altman summed up the result on X: “we made a chip and it is fast.”
Investor reaction focused on whether custom silicon from major AI buyers could pressure Nvidia’s inference dominance. SemiAnalysis argued the fairer comparison is Nvidia’s newer Vera Rubin platform, where Jalapeño still edges ahead on output tokens per megawatt, while total cost per token is roughly even. Part of the savings comes from trading Nvidia’s high margins for Broadcom’s lower margins, reframing the result as a cheaper path to comparable performance rather than outright silicon superiority.