OpenAI unveiled Jalapeño, its first in-house AI inference chip built with Broadcom, on August 25, 2026, at the Hot Chips conference — an ASIC that delivers 1.5–1.9× more tokens per watt than Nvidia’s Blackwell GPU and 2.1–4.1× better interactivity for real-time workloads. Production ramp begins in Q4 2027. For businesses using ChatGPT APIs or OpenAI-powered SaaS tools, faster responses and lower API costs are the expected downstream effect once Jalapeño replaces Nvidia hardware at scale.

What Jalapeño Is and What It Does

Jalapeño is an AI inference chip — a processor designed specifically to run (not train) large language models as fast as possible per unit of power. OpenAI developed it with Broadcom in approximately 16 months, from hiring the team in mid-2024 to taping out the chip in November 2025. The chip is manufactured on TSMC’s N3P process node, uses HBM4 memory at 15.4 TB/s bandwidth, and draws 700W per chip — compared to 900–1,150W for Nvidia’s Rubin GPU.

According to SemiAnalysis analysts Bryan Shan et al., who reviewed the chip’s performance data, Jalapeño reaches approximately 700 tokens per second per user on models including Kimi K2.5 and DeepSeek R1 at concurrency 1. The current A0 engineering samples are already faster than Blackwell on the tested workloads; a B0 stepping, still in fabrication, adds another ~25% improvement in performance per watt. A full rack holds 128 Jalapeño ASICs at approximately 160 kW total power — comparable to an Nvidia GB300 rack.

SemiAnalysis analysts described the chip as follows: “Jalapeño smokes every other chip.” OpenAI also used its own Codex AI to write kernel code during development, which reduced SIMD area by 8% and matrix-engine area by 10%.

What This Means for Businesses Using AI Tools

Businesses that use ChatGPT, GPT APIs, or SaaS products built on OpenAI’s infrastructure should expect two practical changes when Jalapeño reaches production scale in 2027: faster response times and lower API costs. OpenAI’s core constraint today is datacenter power capacity — more tokens of output requires more megawatts. Jalapeño delivers far more tokens per megawatt, which translates directly into cost-per-token reductions as the company replaces Nvidia hardware with its own silicon.

The production schedule matters: most Jalapeño output is scheduled for Q4 2027, not earlier. OpenAI plans to continue using Nvidia and AMD GPUs alongside Jalapeño — this is an addition to its infrastructure, not a full replacement. Businesses that depend on high-volume API access, such as companies running enterprise AI agent workflows, stand to benefit most when per-token costs fall. See how the best AI tools for business currently stack up on cost and speed — those rankings may shift meaningfully by mid-2027.

Important Caveats: What the Benchmarks Do Not Show

The benchmark results published by SemiAnalysis test an “8k1k” workload — 8,000 tokens of input, 1,000 tokens of output. Agentic workloads, which involve multi-turn conversations, long context windows, and tool calls, have not yet been tested on Jalapeño hardware. The results shown are from engineering samples, not production chips. The comparison to Nvidia’s Rubin GPU is, per SemiAnalysis, “incomplete and unfair” because Rubin uses multi-token prediction (MTP) speculative decoding while the Jalapeño results do not.

No specific API price reductions have been announced by OpenAI. The chip’s wider competitive impact on Nvidia is also not settled: Nvidia’s CUDA software ecosystem remains the dominant programming environment for AI workloads, and previous custom chip programs from Meta and Microsoft did not reach comparable scale.

The Competitive Signal for the AI Infrastructure Market

OpenAI’s success completing a competitive inference chip in 16 months — and crediting AI-assisted design (Codex writing kernel code) for part of that speed — is a significant industry benchmark. As SemiAnalysis noted: “OpenAI models like GPT 5.6 Sol, which currently run on NVIDIA GPUs, have been used to design a chip that poses a real threat to the CUDA moat — NVIDIA’s own GPUs are helping usher in their potential successor in real time.” Jensen Huang, Nvidia’s CEO, framed the competitive stakes at Computex 2026: “If you have 1 gigawatt of power, then throughput per watt is revenue.”

For enterprises evaluating AI model providers, chip-level efficiency is now a proxy for long-term pricing power. Models like Anthropic Fable 5 and DeepSeek are already under comparison for enterprise use partly on cost grounds. AI agents for business tasks multiply per-token costs across every automated action — infrastructure efficiency improvements from Jalapeño will flow through to agent economics first.

For Context: OpenAI Infrastructure Coverage on WithO2

This story is part of a series on OpenAI’s infrastructure and model decisions that affect enterprise AI tool users:

Our Take

OpenAI is building infrastructure to cut its own cost base and control the inference layer of the AI market. Businesses that depend on OpenAI APIs stand to benefit from cheaper and faster tokens in 2027 — but the wait is real, the benchmarks are limited to simple workloads, and Nvidia’s training dominance is unaffected. The practical decision for enterprise AI buyers today is unchanged: evaluate providers on current pricing and performance, not on 2027 roadmap assumptions.

Share.

I am a software engineer, I have a passion for working with cutting-edge technologies and staying up-to-date with the latest developments in the field. In my articles, I share my knowledge and insights on a range of topics, including business software, how to set up tools, and the latest trends in the tech industry.

Comments are closed.

Exit mobile version