Not Consulting Huang
Jay, reporting from Aofeisi
The natural order has been completely overturned, folks.
Just now, OpenAI’s in-house chip, “Jalapeño,” crushed NVIDIA’s most powerful Blackwell in benchmark tests…
With a power draw of 700W, it delivers up to 1.9× the throughput per watt, while reducing inference latency to as little as 1/3.6 that of the competing products.
This isn’t official PR. The data comes from the semiconductor research firm SemiAnalysis, which was invited to OpenAI’s labs to conduct on-site testing.
Generally speaking, first-generation chips aren’t competitive.
But SemiAnalysis ultimately reached this conclusion:
OpenAI’s first-generation chip can beat Blackwell in almost every scenario.
Low latency, high throughput, low concurrency… all reaching SOTA-level performance.
Reportedly, Codex played a very important role in the process. Some kernels written by AI even outperformed those written by human engineers.
This is also a form of self-evolution, perhaps.

Interestingly, NVIDIA’s stock price not only failed to plunge, but rose 2.19%…

For one thing, NVIDIA has new products.
On the same day Jalapeño was unveiled, NVIDIA launched the Jetson Orin Nano 2 robotic computer.
For another, it may also have something to do with OpenAI’s stance.
After releasing the Jalapeño benchmark results, OpenAI still explicitly stated that it would not abandon NVIDIA: training would continue to make extensive use of NVIDIA chips, and the two companies also have a financing-guarantee partnership.
The Order Is Reversed! OpenAI’s Chip Crushes Blackwell
In June this year, OpenAI officially announced a partnership with Broadcom to develop this inference-focused chip, named “Jalapeño.” Design work began in mid-2024, and it took only about 16 months to go from assembling the team to tape-out.
On August 25, OpenAI unveiled the self-developed chip at Hot Chips 2026.
At the same time, more hidden details gradually came to light.
SemiAnalysis was invited to OpenAI’s labs with its own InferenceX benchmark suite to conduct on-site testing.
The conclusion: on multiple mainstream open-source models, this first-generation chip outperformed every NVIDIA, AMD, and Google chip within the scope of the tests.

Power efficiency is its biggest highlight.
Jalapeño has a thermal design power (TDP) of just 700W, while Blackwell flagship accelerators have a TDP of 1,200W to 1,400W. SemiAnalysis’s testing showed that its inference throughput per kilowatt reached 1.5 to 1.9 times that of the GB200 NVL72 and GB300 NVL72 rack systems.
Latency was equally impressive.
End-to-end inference latency was 1.7 to 3.6 times lower than the best results recorded by NVIDIA for the GB200 NVL72 and GB300 NVL72. In ultra-low-latency scenarios, it was 2.1 to 4.1 times faster.
At GB300’s fastest output-speed setting, throughput per kilowatt was up to 8.6 to 104.3 times higher.

The test models included three open-source large models: GPT-OSS 120B, DeepSeek R1 670B, and Moonshot AI’s Kimi K2.5 (1 trillion parameters).
Under low-concurrency conditions, GPT-OSS and Kimi K2.5 reached approximately 1,400 tokens per second per user. DeepSeek R1 exceeded 700 tokens per second per user at a concurrency of one.
At the low-concurrency point of 100 tokens per second, Kimi K2.5 performed at more than 9 times the level of the second-best chip. At the same interaction level, GPT-OSS delivered nearly twice the throughput per megawatt of GB200’s highest-throughput point; at a concurrency of one, it was 50 times higher.
Accuracy did not lag behind either: its GSM8K evaluation results were on par with NVIDIA’s chips.
That said, the testing methodology also needs to be clarified.
All of the results above were based on single-token prediction (STP), without speculative decoding or prefill/decode disaggregation. The Blackwell results used for comparison, however, enabled multi-token prediction (MTP). SemiAnalysis said that if compared with the GB300 with multi-token prediction enabled, Jalapeño’s peak power-efficiency advantage would narrow to approximately 1.5×.
The benchmark data was provided by OpenAI. The SemiAnalysis team verified the InferenceX benchmarking process on-site at the lab, but did not run the complete suite or test its preferred AgentX benchmark. The latter covers long contexts and multi-turn conversations, making it closer to real-world production workloads.
At the same time, SemiAnalysis believes that comparing Jalapeño with Blackwell is not entirely fair. Jalapeño’s true benchmark should be the next-generation Vera Rubin, which likewise uses HBM4.
The Rubin system has already begun shipping to customers, while Jalapeño currently exists only as an engineering sample.
But at present, Jalapeño still appears competitive—even against Rubin directly.
Its per-megawatt output throughput for single-token prediction exceeded the multi-token prediction results published by NVIDIA and CoreWeave in July, and was also far ahead of GB200’s multi-token prediction results from 2025.
The two were roughly comparable in output per dollar, while Rubin’s results used speculative decoding. Speculative decoding can further reduce the cost per token by 3 to 5 times. Jalapeño has not yet adopted speculative decoding; once that capability is added, its cost advantage could widen further.

How Did It Do This?
Altman, you… you… that speed is honestly a little terrifying.
After completing tape-out in November 2025 (with the CoWoS packaging design), it took only three months to power on the actual chip, and just nine months to achieve A0 stepping results.
The second-generation B0 stepping has already entered the fab for tape-out, with performance per watt expected to improve by approximately another 25%.

In terms of specifications, the B0 stepping uses a single compute die manufactured on TSMC’s N3P process, paired with an I/O chiplet built on the N3E process. Each die can deliver 13.4 PFLOPS of MXFP4 compute performance, with a TDP of only 700W.
For memory, the chip is equipped with six stacks of HBM4 high-bandwidth memory, providing 216GB per package and bandwidth of 15.4TB/s—the highest among accelerators that have shipped or are about to ship. Samsung is the most likely supplier.

The software stack is equally important. The kernels are written in the Gluon language. Based on Triton, the language retains the SPMD programming model while offering lower-level abstractions.
Reportedly, AI played a very important role in this rapid development process.
During the design phase, Codex and GPT-Astra were used as assistants, reducing the area of the SIMD units by 8% and the matrix engines by 10%.
AI was also responsible for writing every kernel. In some modules, AI-written kernels were 1.5 to 1.8 times faster than those written by human experts.
The iteration speed was astonishing as well.
Within eight days, the parallelism was expanded from TP8 to TP32 across an entire rack; in less than two weeks, throughput more than doubled at certain interaction levels.
There was also a fun Easter egg: OpenAI used Codex to port DOOM to this chip, where it can run at 36 FPS.

Ironically, OpenAI’s GPT 5.6 Sol model, which runs on NVIDIA GPUs, also participated in designing this chip that threatens CUDA’s ecosystem moat.
NVIDIA’s own GPUs are helping bring a potential successor into existence with their own hands.
Having SOTA-level internal AI support may also be one way OpenAI managed to overtake its rivals on the inside track.
When Will Jalapeño Arrive?
So when will this Jalapeño officially enter the arena?
The answer: the engineering samples are ready.
Mass production will gradually ramp up in 2027, with most output concentrated toward the end of the year. The next-stage target is to reach a scale of 100MW.
The single-rack system hosting Jalapeño is called Vindaloo. It is equipped with 128 Jalapeño chips, along with Katsu CPU host racks and Chana switch racks. The two racks together consume approximately 160kW, and up to 16 racks—with 2,048 chips—can be combined into a single scale-up domain.
Deployment will be carried out in partnership with neocloud providers. Reliability data will be collected first, followed by a gradual production ramp-up.
It is worth noting that OpenAI currently does not plan to use Jalapeño to replace its entire existing chip lineup. It will continue working with NVIDIA and other computing partners.
At the same time, OpenAI will continue developing second- and third-generation chips.
One More Thing
However, who will successfully lead this army of peppers remains a question mark.
Freshly surfaced reports indicate that Chris Malone, head of OpenAI’s data centers, left the company last week.

The Wall Street Journal was the first to report the news on August 25, followed by Bloomberg, which confirmed it through an OpenAI spokesperson. After Malone’s departure, there will be no successor; his responsibilities have been distributed among several leaders.
He joined OpenAI in March 2025, shortly after the company announced the Stargate project with Oracle and SoftBank, and was responsible for building data centers to meet the company’s rapidly growing compute needs.
Stargate got off to a rocky start. OpenAI temporarily shifted to renting chips from cloud providers, and has recently begun returning to self-built infrastructure—specifically, leasing entire facilities rather than individual chips. However, according to people familiar with the matter, Malone was not the person leading that effort.
Malone is the fourth executive to leave during this wave of departures.
Previously, Chief Revenue Officer Denise Dresser, Chief Operating Officer Brad Lightcap, and CEO Altman’s deputy Fidji Simo had all left. TheNextWeb reported that since April, seven senior executives have either departed or changed roles.
Such a concentrated wave of C-suite departures in the months before a public listing is rare, even by the standards of the eve of an IPO.
Jalapeño seems to have no one left to take over…
How about asking Huang to send someone over to put out the fire?
(Just kidding.)
References: