In-depth | OpenAI's First-Generation In-House Chip Shines, Unveiling a New Approach to Computing Power Beyond Expectations

09/10 2026 409

Foreword:

A chip named Jalapeño has propelled OpenAI onto a different playing field. The first-generation in-house AI chip has achieved top-tier performance in real-world tests, with unit power throughput and latency metrics placing it in the first echelon.

Interestingly, OpenAI has also changed the metric for measuring computing power: less focus on peak performance per chip, more on how much useful intelligence can be generated per megawatt of electricity.

Author | Fang Wensan

Image Source | Internet

First-Generation Success: OpenAI Changes the Metric

Recently, OpenAI released the first batch of real-world performance data for its first in-house inference chip, Jalapeño.

Across three public models, Jalapeño's peak performance per unit of power consumption is approximately 1.5 to 1.9 times that of competing systems, with end-to-end latency reduced by about 1.7 to 3.6 times. Under high-interaction workloads, performance advantages reach 2.1 to 4.1 times. Jalapeño has a rated power consumption of 700W, with OpenAI stating that sustained power consumption under test workloads does not exceed 550W.

Current testing primarily targets fixed-length inference scenarios, with OpenAI's official appendix using configurations such as 8k input and 1k output. SemiAnalysis cautions that comparisons between Jalapeño and GB200/GB300 do not cover all competitive realities, as NVIDIA's Vera Rubin has entered a new product cycle. More complex long-context and multi-round agent workloads have also not been fully validated through AgentX.

Thus, the conclusions Jalapeño can currently support are clear: it has proven OpenAI's ability to create a competitive inference ASIC.

The truly unexpected aspect lies in the chip's internal design philosophy. LLM inference does not involve continuous identical computations. Prompt pre-filling is computationally intensive, token-by-token generation is more memory bandwidth-dependent, and inter-chip communication introduces latency.

OpenAI reorganizes computation, memory, KV Cache, and networking around these real-world workloads, minimizing data movement and keeping model states closer to computation. Chips, memory, networking, scheduling software, and server racks are thus designed as a single machine.

Inference Takes the Lead: Every Token Comes with a Cost

OpenAI's decision to focus its first chip on inference is telling. Training cutting-edge models still requires immense general-purpose computing power, mature software ecosystems, and large-scale clusters—areas where GPUs, particularly NVIDIA's, maintain a strong advantage.

OpenAI has also made clear that it will continue to use NVIDIA and other partners' accelerators at scale for training and inference.

The economics of inference, however, are entirely different. Once a model is trained, every ChatGPT response, API call, Codex task, and multi-step agent operation generates new computational consumption.

In the age of agents, latency begins to have a "compounding effect." A slight delay in one step can accumulate over dozens of serial steps, resulting in a prolonged silence for the user.

OpenAI thus focuses testing on unit power throughput, per-user token speed, and end-to-end latency, rather than merely showcasing theoretical chip performance.

As model companies scale their user bases, inference transforms from a technical back-end consideration into a long-term variable in profit statements. Generating more tokens per watt, supporting more users per server, and consuming fewer resources per request can all alter service costs when aggregated across millions or even billions of calls.

OpenAI CFO Sarah Friar has articulated this approach clearly in discussing computing power strategy: select different hardware based on workloads, use high-end systems where capabilities matter most, pursue efficiency where scale and cost are more sensitive, and ultimately measure success by "how much useful intelligence is generated per dollar."

Thus, Jalapeño is not merely a semiconductor project; it is more akin to a "gross margin machine" tailored to OpenAI's business model.

In-House Chips Don't End Procurement: Computing Power Becomes an "Investment Portfolio"

Seeing OpenAI develop chips might lead one to assume the goal is to reduce NVIDIA dependence. This is only partially true. OpenAI is developing in-house ASICs while also aggressively expanding external procurement.

In September 2025, OpenAI and NVIDIA announced plans to deploy at least 10GW of NVIDIA systems, with the first 1GW using the Vera Rubin platform. In October 2025, OpenAI reached a 6GW GPU cooperation agreement with AMD, with the first 1GW of MI450 series deployments slated to begin in the second half of 2026.

Meanwhile, in 2025, OpenAI and Broadcom announced a 10GW cooperation plan for in-house accelerators, targeting deployments from the second half of 2026 through 2029. By September of this year, Broadcom disclosed that OpenAI-related committed capacity in visible projects through 2028 exceeded 5GW.

These figures cannot be mechanically added, as different agreements may overlap in timing, data centers, supply methods, and workload types. However, they paint a clear picture: OpenAI is abandoning the mindset of seeking a single "best chip."

For training cutting-edge models, the strongest general-purpose GPUs can be used; for high-frequency, large-scale inference, OpenAI's own ASICs can be gradually adopted; for workloads requiring extremely low latency, Cerebras can be integrated; different cloud platforms continue to provide elastic capacity; and data centers and energy are secured in advance through infrastructure initiatives like Stargate.

In the past, when alternatives were scarce, chip suppliers' pricing essentially set the baseline for computing power costs. With a truly functional in-house development route, OpenAI now holds additional leverage at the negotiating table, even as it continues to procure large volumes from NVIDIA.

Thus, Jalapeño's strategic value need not rest on "replacing NVIDIA" as a grand premise; it need only become a credible option to alter procurement relationships.

Nine-Month Tape-Out: AI Begins to Transform Chip Development

Another easily overlooked aspect of Jalapeño is the change it represents: AI is now involved in developing the next generation of AI chips.

OpenAI states that Jalapeño's joint development cycle from initial design to manufacturing tape-out was approximately nine months, with models used to explore different implementation options, shorten design and verification cycles, and optimize some arithmetic circuits. After the chip entered software development, AI participated in kernel optimization, task mapping, and scheduling.

On selected GPT-OSS attention and MoE modules, AI-generated implementations ran 1.5 to 1.8 times faster than code written by human experts.

While this figure applies only to selected modules and cannot be extrapolated to overall model performance, it reveals a compelling new cycle: models help design chips, chips enable faster execution of the next generation of models, and new models then participate in optimizing the next generation of chips.

OpenAI is attempting to compress precisely this time lag. Jalapeño has not yet proven that this nine-month rhythm can be sustained long-term, but OpenAI has already disclosed that the second-generation product is in advanced development, with the third generation taking shape.

Once this R&D loop is established, new sources of competitive advantage will emerge for in-house ASICs: knowing sooner than anyone else what the next generation of models will require.

OpenAI has confirmed that Gen 2 is in advanced development and Gen 3 has begun. This means NVIDIA's true concern should not be how many tokens Jalapeño A0 can process in this generation, but whether OpenAI can accelerate this loop with each subsequent generation.

In-House Doesn't Mean In-House Manufacturing: Industry Chain Profits Are Reshuffled

Jalapeño also reminds the market that "in-house chip development" is often romanticized.

OpenAI handles core architectural design, Broadcom manages chip implementation, networking, and connectivity technologies, Celestica oversees board-level and rack-level system industrialization, and TSMC manufactures the chips.

OpenAI gains stronger chip control but does not become a wafer foundry; dependencies are simply reconfigured.

Previously concentrated among GPU vendors, profits are now dispersing into ASIC design services, advanced process nodes, HBM, advanced packaging, networking, server manufacturing, and data center power. One of the companies most directly benefiting from this shift is Broadcom.

In the third quarter of fiscal year 2026 (ending August 2), Broadcom's AI semiconductor revenue reached $16.7 billion, up 221% year-over-year; the company expects this figure to rise further to $21.7 billion in the fourth quarter.

These numbers reveal an industrial shift larger than "who challenges NVIDIA": the AI computing power market is stratifying.

General-purpose GPUs still dominate cutting-edge training, ecosystems, and highly flexible workloads; model vendors with hyperscale, stable demand are increasingly motivated to solidify part of their computation into custom ASICs.

Companies like Broadcom and Marvell, capable of providing custom chips, interconnectivity, and system-level capabilities, are thus positioned at new traffic entry points.

For NVIDIA, this does not yet spell collapse. OpenAI's simultaneous development of Jalapeño and continued planning for massive NVIDIA infrastructure speaks volumes. What is truly shifting is the era's imagination of "one chip to rule all AI computing."

Conclusion:

Jalapeño does not conclude the AI chip war; it merely advances the rules by one notch. OpenAI now simultaneously holds both model and hardware ends, enabling algorithmic requirements to directly inform chip design and making every watt of electricity a participant in commercial competition.

Future disparities among model companies will also reside in those increasingly bespoke, increasingly specialized silicon wafers.

OpenAI: "Jalapeño's Initial Test Results Demonstrate Industry-Leading AI Inference Speed and Efficiency," OpenAI: "OpenAI and NVIDIA Announce Strategic Partnership to Deploy 10 Gigawatts of AI Data Centers Using NVIDIA Systems," Reuters: "OpenAI Builds First Chip with Broadcom and TSMC, Scales Back Foundry Ambitions"

Solemnly declare: the copyright of this article belongs to the original author. The reprinted article is only for the purpose of spreading more information. If the author's information is marked incorrectly, please contact us immediately to modify or delete it. Thank you.