On August 25, OpenAI dropped the first official benchmarks for its in-house inference chip, Jalapeño. According to the company, Jalapeño delivered 1.5–1.9x more compute per watt and cut latency by 1.7–3.6x versus systems built on Nvidia's GB200 and GB300. ForkLog first reported on the release.

What OpenAI Says

Jalapeño is OpenAI’s debut custom chip for inference workloads. The company claims the test results show a major leap in efficiency: more compute for every watt consumed, and faster response times. Their internal benchmarks highlight a combo of higher throughput and lower latency, all within a single architecture.

How It Stacks Up Against Nvidia Blackwell

The comparison was run against Nvidia Blackwell-based systems—specifically, the GB200 and GB300 configs. In these tests, Jalapeño pulled off a 1.5–1.9x edge in work per watt, and slashed latency by 1.7–3.6x. All these numbers come straight from OpenAI’s own measurements.

OpenAI is pitching these results as a big step forward in inference efficiency. But they haven’t shared details on their testing methodology or hardware setups; the comparison is strictly about energy and latency metrics versus GB200 and GB300 systems.