Blog

OpenAI's 'Jalapeno' Chip: Custom Silicon to Take on Nvidia

tech Aug 26, 2026 4 min read By Pyae Phyo Kyaw

OpenAI's first custom AI chip, Jalapeno — built with Broadcom — is a purpose-built inference accelerator that claims up to 1.9x the work per watt of Nvidia's GB200 and 1.7x–3.6x lower latency [1][2]. Unveiled in June and detailed at Hot Chips in August, it is the first step in OpenAI's plan to cut dependence on Nvidia silicon, with gigawatt-scale data-center deployment starting by end-2026 [1][2].

Story at a glance

EVENT — OpenAI and Broadcom unveil JalapenoFirst custom chip; ~9 months RTL to tape-outIMPACT — Up to 1.9x work per watt vs GB200Lower latency, up to 50% lower inference costHISTORICAL PARALLEL — Google's TPU journeyCustom silicon grew from niche to strategicFUTURE OUTLOOK — Deployment from end-2026Gen 2 in design; still buys Nvidia for training
Figure 1: OpenAI's Jalapeno chip at a glance — the launch, the performance claims, the TPU parallel, and the roadmap.

What: an inference-only ASIC

Jalapeno is an application-specific integrated circuit for large-language-model inference — running trained models — not for training [1][2]. Broadcom handled silicon implementation and networking; Celestica the systems integration [1]. It went from initial RTL to tape-out in about nine months, one of the fastest advanced-chip development cycles on record [1].

Why it matters: the economics of serving AI

Inference is where AI costs land. OpenAI claims Jalapeno could cut inference costs by up to 50% versus current GPUs, while delivering better performance per watt [1][2]. If true, it changes the cost structure of every ChatGPT request and pressures Nvidia's pricing power [1][2].

Who: OpenAI, Broadcom, and the ecosystem

The chip pairs OpenAI with Broadcom, whose CEO called it "just the beginning" [2]. OpenAI used its own models to help design the chip, and it runs models from other providers too [2]. Microsoft and other partners will host the first deployments [1][2].

When: from concept to deployment

Table 1: Jalapeno timeline. Sources: ServeTheHome, Khaleej Times [1][2].
DateMilestone
Late 2024Architecture concept
Late 2025Tape-out (~9 months from RTL)
Early 2026Codex running on the chip
Jun 2026Chip unveiled
End-2026Gigawatt-scale data-center deployment

Where: gigawatt-scale data centers

Jalapeno is designed for gigawatt-scale deployments with Microsoft and partners starting in late 2026, with a broader ramp in 2027–2028 [1][2]. A single global system spans 2,048 chips reaching 27 exaflops [1].

Which: the numbers that matter

Table 2: Jalapeno vs Nvidia claims. Sources: ServeTheHome [1].
MetricJalapeno claim
Work per watt (peak)1.5x–1.9x vs GB200/GB300
End-to-end latency1.7x–3.6x lower
Agentic workloads2.1x–4.1x higher performance
Package power700W (vs GB300's 1.4kW)
Memory15.4 TB/s HBM4, 216 GiB
Inference costUp to 50% lower (claimed)

How: designed around the inference bottleneck

Inference has two phases: compute-heavy prefill and memory-bandwidth-heavy decode. Jalapeno balances both, keeps model state and cache local to avoid network delays, and gates idle blocks per phase [1]. OpenAI says AI-generated code it produced for the design runs up to 1.8x faster than expert-written versions [1].

What next: a multi-generation platform

Historical parallel

The playbook echoes Google's TPU: what began in 2015 as a complement to Nvidia GPUs for inference became a strategic, multi-generation platform that now anchors Google's AI infrastructure [1][2]. OpenAI is following the same path — custom silicon first for inference, then deeper [1].

Future outlook

  • Cost advantage: If the 50% cost claim holds at scale, OpenAI can serve more users per dollar and pressure rivals [1][2].
  • Nvidia stays for training: Jalapeno is inference-only; OpenAI says it will keep buying Nvidia chips for training [1].
  • Roadmap risk: Gen 2 is in design, but scaling a brand-new chip into gigawatt data centers is hard — delays are the key risk [1].

What to watch: real-world performance at Microsoft data centers, Gen 2 tape-out, and whether Nvidia's upcoming Vera Rubin reasserts the performance lead [1][2].

References

  1. ServeTheHome — OpenAI Jalapeno Custom AI ASIC at Hot Chips 2026
  2. Khaleej Times — OpenAI unveils first custom-designed AI chip 'Jalapeno'
  3. Mobile World Live — OpenAI claims Jalapeno chip outperforms Nvidia GB300

Disclaimer

This content is for educational purposes only. Figures are as of 26 August 2026 and may be revised. Always verify current information before relying on it.