Published: August 17, 2026
OpenAI launched Ultrafast on August 13, 2026 — a new API service tier that runs GPT-5.6 Sol up to 14× faster than Standard processing, generating up to 750 output tokens per second on Cerebras hardware. OpenAI released Ultrafast as a limited preview to a select group of customers, with no published price and no general-availability date. The tier covers GPT-5.6 Sol only and launches first in the OpenAI API, targeting businesses that run real-time workloads such as coding, customer support, financial research, and commerce.
What Ultrafast Changes for Business Workloads
Ultrafast removes the trade-off between model intelligence and response speed: businesses previously had to pick a smaller model to get real-time latency, and Ultrafast delivers OpenAI’s most intelligent model at 750 tokens per second instead. OpenAI names 4 target scenarios: incident response, where the model analyzes logs and code changes while an outage is still unfolding; financial research, assessing transactions as conditions change; voice support, resolving multi-step customer issues without pauses in the conversation; and commerce, answering product and checkout questions before a shopper abandons the cart. According to Courtland Lykins, Product Lead for Voice AI at Podium, an early tester: “The speed completely changes the call experience for the more complex work.” Alex Wang of Applied AI at Rogo said Ultrafast “makes complex financial research feel like a real-time interaction.”
The Cerebras Partnership Behind the Speed
Cerebras powers Ultrafast’s inference, extending a partnership that The Decoder reports was signed earlier in 2026 at a value of $10 billion. OpenAI’s announcement states the collaboration brings “ultra-low-latency inference” to its platform but discloses no chip models, cluster sizes, or capacity figures. OpenAI also uses Ultrafast internally for 2 workflows: incident response, where engineers validate fixes “in a fraction of the time,” and research, where overnight experiment-review loops tighten into multiple same-day iterations.
Pricing: A Three-Tier Speed Ladder Takes Shape
OpenAI published no Ultrafast pricing at launch. The existing API Fast Mode runs GPT-5.6 Sol up to 2.5× faster than Standard at roughly double the price, and The Decoder notes Ultrafast creates a third tier that “turns inference speed into its own product.” All speed figures carry an “up to” qualifier, and OpenAI states access will expand “as capacity grows” via an enterprise interest form. Teams comparing the best AI tools for business should treat inference speed as a priced dimension alongside capability from here on.
For Context
WithO2 has covered GPT-5.6 Sol since its debut: our report on the GPT-5.6 Sol public launch tracks the model’s rollout and already notes the Cerebras 750 tokens-per-second work, and our original GPT-5.6 Sol, Terra, and Luna launch coverage explains how the three variants split OpenAI’s lineup.
Our Take
Ultrafast is a pricing story wearing a benchmark headline. OpenAI now sells the same model at three speeds — Standard, Fast at roughly 2× the price, and a preview tier that is “likely pricier” still, per The Decoder — which mirrors how cloud providers monetize compute performance classes. For SMBs, the practical question is which workflows justify the premium: voice agents and checkout flows feel speed directly; batch content and analysis jobs do not. Readers weighing execution-capable AI for those real-time flows can compare options in our guide to the best AI agents for business tasks.