News

/
News

OpenAI has released the actual test results of its first self-developed AI chip

On August 26, 2026, OpenAI released the first batch of test results for its self-developed AI inference chip, Jalapeño.


In tests with models such as GPT-OSS 120B, DeepSeek R1 670B, and Kimi K2.5 1T, the peak per-watt inference 

performance of Jalapeño reached approximately 1.5 to 1.9 times that of NVIDIA's GB200 and GB300 comparison 

systems. At the same time, the end-to-end latency was reduced to approximately 1/1.7 to 1/3.6.


In the GPT-OSS 120B test, the peak throughput of Jalapeño reached 85,448 mixed tokens/s per kilowatt, which was

 1.9 times higher than NVIDIA's GB200 at 44,960. The end-to-end latency was 1.03 seconds and 1.80 seconds 

respectively. The rated power consumption of Jalapeño was 700W, while that of GB200 was 1,200W.


In the test with the larger parameter scale of DeepSeek R1 670B, Jalapeño was compared with NVIDIA's GB300.


The peak per-kilowatt throughput of Jalapeño was 19,641 mixed tokens/s, while GB300 was 11,781. Jalapeño was 

approximately 1.7 times higher than the latter; the end-to-end latency was 1.65 seconds, while GB300 was 5.99 

seconds, reducing by approximately 3.6 times.


In the Kimi K2.5 1T test, the peak per-kilowatt throughput of Jalapeño was 18,195 mixed tokens/s, while GB300 was

 11,862. The former was approximately 1.5 times higher; the end-to-end latency was reduced from 5.31 seconds of

 GB300 to 1.56 seconds, approximately 3.4 times lower.


OpenAI stated that all three sets of tests were based on the public inference benchmark InferenceX of SemiAnalysis,

 mainly measuring how different systems can complete AI inference work with a unit of power under the same user

 experience and latency requirements. OpenAI normalized each accelerator according to the public rated power 

consumption, with Jalapeño's rated power consumption being 700W, but the actual measured power consumption

 during this test did not exceed 550W.


Jalapeño is an ASIC specifically designed for large language model inference. This test mainly compared inference
throughput, power efficiency, and latency, and did not cover dimensions such as cutting-edge model training, 

general GPU computing capabilities, and software ecosystems. Therefore, a more accurate statement is that 

Jalapeño exceeded NVIDIA's GB200 and GB300 in some of OpenAI's large model inference tests.


As the usage of ChatGPT, Codex, API, and AI Agent increases, inference has become an important component of 

OpenAI's computing expenditure. OpenAI plans to deploy Jalapeño to its own computing infrastructure by the end

 of 2026: For OpenAI, the significance of the self-developed chip is not just to compete with NVIDIA for performance

 leadership. It is gradually integrating models, inference software, chips, networks, and data centers into the same 

technical system, reducing the unit AI inference cost through software and hardware collaboration.


Jalapeño was officially released in June 2026 and is OpenAI's first self-developed AI accelerator. OpenAI is responsible 

for designing the chip architecture from scratch, Broadcom participates in the chip implementation, network and 

connection technologies, and Celestica is in charge of circuit boards, racks and system integration. OpenAI stated 

that the chip was designed and fabricated in just 9 months, and during the process, OpenAI's models were extensively

 used to assist in chip design and optimization. Unlike NVIDIA's GPUs sold in the general AI computing market, the

 core goal of Jalapeño is to reduce OpenAI's continuously rising model inference costs.


Jalapeño is also the first product of OpenAI's multi-generation self-developed chip roadmap. The second-generation 

chip has entered the in-depth development stage, and the third-generation product has already begun planning. 

Jalapeño optimizes specifically for large model memory movement, KV Cache, network communication, and Prefill,
 Decode, etc. in the inference stage, hoping to serve more users under the same power conditions. Previously, 

OpenAI also disclosed that the multi-generation computing platform where Jalapeño is located will expand to 

gigawatt-level deployment scale in collaboration with data center partners. OpenAI also clearly stated that it will 

continue to use accelerators from NVIDIA and other partners, and select appropriate hardware based on different 

training and inference loads.