OpenAI's 'Jalapeno' Chip: Custom Silicon to Take on Nvidia
OpenAI's first custom AI chip, Jalapeno — built with Broadcom — is a purpose-built inference accelerator that claims up to 1.9x the work per watt of Nvidia's GB200 and 1.7x–3.6x lower latency [1][2]. Unveiled in June and detailed at Hot Chips in August, it is the first step in OpenAI's plan to cut dependence on Nvidia silicon, with gigawatt-scale data-center deployment starting by end-2026 [1][2].
Story at a glance
What: an inference-only ASIC
Jalapeno is an application-specific integrated circuit for large-language-model inference — running trained models — not for training [1][2]. Broadcom handled silicon implementation and networking; Celestica the systems integration [1]. It went from initial RTL to tape-out in about nine months, one of the fastest advanced-chip development cycles on record [1].
Why it matters: the economics of serving AI
Inference is where AI costs land. OpenAI claims Jalapeno could cut inference costs by up to 50% versus current GPUs, while delivering better performance per watt [1][2]. If true, it changes the cost structure of every ChatGPT request and pressures Nvidia's pricing power [1][2].
Who: OpenAI, Broadcom, and the ecosystem
The chip pairs OpenAI with Broadcom, whose CEO called it "just the beginning" [2]. OpenAI used its own models to help design the chip, and it runs models from other providers too [2]. Microsoft and other partners will host the first deployments [1][2].
When: from concept to deployment
| Date | Milestone |
|---|---|
| Late 2024 | Architecture concept |
| Late 2025 | Tape-out (~9 months from RTL) |
| Early 2026 | Codex running on the chip |
| Jun 2026 | Chip unveiled |
| End-2026 | Gigawatt-scale data-center deployment |
Where: gigawatt-scale data centers
Jalapeno is designed for gigawatt-scale deployments with Microsoft and partners starting in late 2026, with a broader ramp in 2027–2028 [1][2]. A single global system spans 2,048 chips reaching 27 exaflops [1].
Which: the numbers that matter
| Metric | Jalapeno claim |
|---|---|
| Work per watt (peak) | 1.5x–1.9x vs GB200/GB300 |
| End-to-end latency | 1.7x–3.6x lower |
| Agentic workloads | 2.1x–4.1x higher performance |
| Package power | 700W (vs GB300's 1.4kW) |
| Memory | 15.4 TB/s HBM4, 216 GiB |
| Inference cost | Up to 50% lower (claimed) |
How: designed around the inference bottleneck
Inference has two phases: compute-heavy prefill and memory-bandwidth-heavy decode. Jalapeno balances both, keeps model state and cache local to avoid network delays, and gates idle blocks per phase [1]. OpenAI says AI-generated code it produced for the design runs up to 1.8x faster than expert-written versions [1].
What next: a multi-generation platform
Historical parallel
The playbook echoes Google's TPU: what began in 2015 as a complement to Nvidia GPUs for inference became a strategic, multi-generation platform that now anchors Google's AI infrastructure [1][2]. OpenAI is following the same path — custom silicon first for inference, then deeper [1].
Future outlook
- Cost advantage: If the 50% cost claim holds at scale, OpenAI can serve more users per dollar and pressure rivals [1][2].
- Nvidia stays for training: Jalapeno is inference-only; OpenAI says it will keep buying Nvidia chips for training [1].
- Roadmap risk: Gen 2 is in design, but scaling a brand-new chip into gigawatt data centers is hard — delays are the key risk [1].
What to watch: real-world performance at Microsoft data centers, Gen 2 tape-out, and whether Nvidia's upcoming Vera Rubin reasserts the performance lead [1][2].
References
- ServeTheHome — OpenAI Jalapeno Custom AI ASIC at Hot Chips 2026
- Khaleej Times — OpenAI unveils first custom-designed AI chip 'Jalapeno'
- Mobile World Live — OpenAI claims Jalapeno chip outperforms Nvidia GB300
Disclaimer
This content is for educational purposes only. Figures are as of 26 August 2026 and may be revised. Always verify current information before relying on it.