Faster than Standard
GPT-5.6 Sol on Ultrafast vs. the same model on Standard processing.
OpenAI · August 13, 2026 · “Previewing Ultrafast mode: GPT-5.6 Sol at up to 14× the speed”
Ultrafast mode runs GPT-5.6 Sol — OpenAI’s most intelligent model — up to 14× faster than Standard processing, at up to 750 output tokens per second. Powered by the largest chip ever built: the Cerebras Wafer-Scale Engine 3.
01 — The race
Both lanes stream the same completion. Ultrafast runs at up to 750 output tokens/s; Standard crawls along at 1/14 of that. This is the real ratio, played live on your screen.
Until now
Real-time speed meant choosing a smaller or more specialized model. Fast or smart — pick one.
Ultrafast
Removes the trade-off. The most intelligent model, at real-time speed. More useful work per second.
02 — The numbers
Scroll them into view — each one counts up and banks points into your Ultrafast IQ. Every figure is from the announcement or Cerebras’ WSE-3 spec sheet.
GPT-5.6 Sol on Ultrafast vs. the same model on Standard processing.
Peak generation speed — a full paragraph in the blink of a cursor.
Etched onto a single piece of silicon — the WSE-3 is the largest chip ever built.
52× more compute cores than the largest GPU — all on one wafer.
880× more on-chip memory than an H100-class GPU. The model lives on the silicon.
21 petabytes per second of on-wafer memory bandwidth — no off-chip bottleneck.
125 petaFLOPS of AI compute on a 5nm TSMC process.
57× larger than the biggest GPU die (814 mm²). Plus 214 Pb/s of fabric bandwidth.
03 — Scale check
A conventional GPU is a small die diced from a wafer. Cerebras skipped the dicing: the WSE-3 is the wafer. Drag the slider — or use ←/→ arrow keys — to overlay them at true relative area.
Drawn to scale by area: the wafer’s face holds ~57 H100-class dies — with room to spare.
04 — Anatomy of a monster
Keep scrolling — the wafer zooms from “dinner plate” to “city seen from orbit,” unlocking each subsystem as you descend.
Subsystem 1 / 6
Every core is a tiny compute engine with its own slice of memory. 52× more cores than the largest GPU — and they all talk to each other without ever leaving the wafer.
05 — Where speed wins
These are the five real scenarios from the announcement. Catch the falling tokens in the glowing lane to unlock each one — mouse, touch, or A/D and ←/→.
Scenario 01
Analyze logs, recent code changes, and engineer reports to find the likely cause — and help prepare a fix while the outage is still unfolding.
Why speed wins: the evidence is changing right now. A fix that arrives after the post-mortem isn’t a fix.
Scenario 02
Analyze market signals, assess transactions, and spot suspicious activity while conditions are still changing.
Why speed wins: markets and fraudsters don’t wait for a batch job to finish.
Scenario 03
Resolve complex issues in real time without interrupting the conversation — even across multiple steps or systems.
Why speed wins: a pause in voice is a hang-up. Latency is the user experience.
Scenario 04
Answer product questions, check inventory, personalize recommendations, and resolve checkout issues before hesitation becomes an abandoned cart.
Why speed wins: hesitation has a half-life measured in seconds.
Scenario 05
Turn overnight research runs into interactive working sessions — multiple iterations inside a single workday.
Why speed wins: when the loop fits in a day, the day produces the breakthrough.
06 — Early voices
A select group of customers is testing Ultrafast in the limited preview — and OpenAI runs it internally, too.
“The increase in speed brought by Cerebras is impressive. It enables different ways of using the models, and makes it practical for developers to work in a more focused and productive way alongside them.”John Crepezzi · AI Assistants, Jane Street
Reading logs, analyzing traces, synthesizing conversations, and preparing or validating fixes — while the evidence is still changing.
Tightening the overnight-batch loop into multiple same-day iterations. The experiment cycle becomes a conversation.
07 — Prove it
Six questions. Every answer is a number you just scrolled past. Streaks multiply your points.
Ready? Wrong answers cost nothing but pride. Streaks of 2+ earn a ×1.5 multiplier.
08 — Get access
Ultrafast mode is available today as a limited preview for a select group of customers in the OpenAI API, expanding as capacity grows. OpenAI has a sign-up form for access updates.
Demo form — nothing leaves this page. The real sign-up lives at openai.com/index/previewing-ultrafast.