ULTRAFAST×CEREBRAS

OpenAI · August 13, 2026 · “Previewing Ultrafast mode: GPT-5.6 Sol at up to 14× the speed”

The speed limit
just got deleted.

Ultrafast mode runs GPT-5.6 Sol — OpenAI’s most intelligent model — up to 14× faster than Standard processing, at up to 750 output tokens per second. Powered by the largest chip ever built: the Cerebras Wafer-Scale Engine 3.

14× faster 750 tok/s 4T transistors 900K cores

01 — The race

Watch 750 tokens a second humiliate the old normal.

Both lanes stream the same completion. Ultrafast runs at up to 750 output tokens/s; Standard crawls along at 1/14 of that. This is the real ratio, played live on your screen.

⚡ ULTRAFAST · GPT-5.6 Sol 0 tok/s
0%
0 tokens
🐢 STANDARD · same model, 1/14 the pace 0 tok/s
0%
0 tokens

Until now

Real-time speed meant choosing a smaller or more specialized model. Fast or smart — pick one.

Ultrafast

Removes the trade-off. The most intelligent model, at real-time speed. More useful work per second.

02 — The numbers

Stats that hit like a tracer round.

Scroll them into view — each one counts up and banks points into your Ultrafast IQ. Every figure is from the announcement or Cerebras’ WSE-3 spec sheet.

0×

Faster than Standard

GPT-5.6 Sol on Ultrafast vs. the same model on Standard processing.

0tok/s

Output tokens / second

Peak generation speed — a full paragraph in the blink of a cursor.

0trillion

Transistors

Etched onto a single piece of silicon — the WSE-3 is the largest chip ever built.

0cores

AI-optimized cores

52× more compute cores than the largest GPU — all on one wafer.

0GB

On-chip SRAM

880× more on-chip memory than an H100-class GPU. The model lives on the silicon.

0PB/s

Memory bandwidth

21 petabytes per second of on-wafer memory bandwidth — no off-chip bottleneck.

0PFLOPS

AI compute

125 petaFLOPS of AI compute on a 5nm TSMC process.

0mm²

Of silicon

57× larger than the biggest GPU die (814 mm²). Plus 214 Pb/s of fabric bandwidth.

03 — Scale check

One wafer. 57 GPUs’ worth of silicon.

A conventional GPU is a small die diced from a wafer. Cerebras skipped the dicing: the WSE-3 is the wafer. Drag the slider — or use ←/→ arrow keys — to overlay them at true relative area.

WSE-3 the entire wafer is one chip GPU
CEREBRAS WSE-3 · 46,225 mm²
LARGEST GPU DIE · 814 mm²
57×larger than the largest GPU die
52×more compute cores
880×more on-chip memory
5nmTSMC process

Drawn to scale by area: the wafer’s face holds ~57 H100-class dies — with room to spare.

04 — Anatomy of a monster

Scroll through 4 trillion transistors.

Keep scrolling — the wafer zooms from “dinner plate” to “city seen from orbit,” unlocking each subsystem as you descend.

ZOOM ×1.0

Subsystem 1 / 6

900,000 AI-optimized cores

Every core is a tiny compute engine with its own slice of memory. 52× more cores than the largest GPU — and they all talk to each other without ever leaving the wafer.

05 — Where speed wins

Five arenas. One catch.

These are the five real scenarios from the announcement. Catch the falling tokens in the glowing lane to unlock each one — mouse, touch, or A/D and /.

⚡ Token Catch

Tokens fall into five lanes — one per real-world scenario. Catch them in the glowing lane to unlock that scenario’s story. Wrong-lane catches still score, but unlock nothing.

A left · D right · or drag / move your mouse

0 pts ❤❤❤ 0 / 5 scenarios
🔒

Scenario 01

Incident response & reliability

Analyze logs, recent code changes, and engineer reports to find the likely cause — and help prepare a fix while the outage is still unfolding.

Why speed wins: the evidence is changing right now. A fix that arrives after the post-mortem isn’t a fix.

🔒

Scenario 02

Financial research & security

Analyze market signals, assess transactions, and spot suspicious activity while conditions are still changing.

Why speed wins: markets and fraudsters don’t wait for a batch job to finish.

🔒

Scenario 03

Customer support & voice

Resolve complex issues in real time without interrupting the conversation — even across multiple steps or systems.

Why speed wins: a pause in voice is a hang-up. Latency is the user experience.

🔒

Scenario 04

Commerce

Answer product questions, check inventory, personalize recommendations, and resolve checkout issues before hesitation becomes an abandoned cart.

Why speed wins: hesitation has a half-life measured in seconds.

🔒

Scenario 05

Live research & experimentation

Turn overnight research runs into interactive working sessions — multiple iterations inside a single workday.

Why speed wins: when the loop fits in a day, the day produces the breakthrough.

06 — Early voices

Already in the fast lane.

A select group of customers is testing Ultrafast in the limited preview — and OpenAI runs it internally, too.

“The increase in speed brought by Cerebras is impressive. It enables different ways of using the models, and makes it practical for developers to work in a more focused and productive way alongside them.”
John Crepezzi · AI Assistants, Jane Street
Jane Street Podium Basis Rogo

🛠 Inside OpenAI: incident response

Reading logs, analyzing traces, synthesizing conversations, and preparing or validating fixes — while the evidence is still changing.

🔬 Inside OpenAI: research

Tightening the overnight-batch loop into multiple same-day iterations. The experiment cycle becomes a conversation.

07 — Prove it

The Ultrafast IQ check.

Six questions. Every answer is a number you just scrolled past. Streaks multiply your points.

Ready? Wrong answers cost nothing but pride. Streaks of 2+ earn a ×1.5 multiplier.

08 — Get access

Limited preview. Expanding fast.

Ultrafast mode is available today as a limited preview for a select group of customers in the OpenAI API, expanding as capacity grows. OpenAI has a sign-up form for access updates.

Demo form — nothing leaves this page. The real sign-up lives at openai.com/index/previewing-ultrafast.

View more demos Get $10 off Kimi K3