OpenAI Ultrafast Mode: Pricing, Speed, and How to Access It
OpenAI's Ultrafast inference mode promises 300 tokens per second on Cerebras hardware, but it's locked to the $500/month Pro-500 plan at 6x cost.

What is OpenAI’s Ultrafast mode?
Ultrafast is a new inference tier OpenAI announced at its dev day event that generates responses at roughly 300 tokens per second, about eight times faster than standard ChatGPT output. It reportedly runs on Cerebras chips rather than OpenAI’s usual inference stack, which is how it hits that kind of throughput. The catch: Sam Altman said Ultrafast will cost 6x more than standard inference, and access is restricted to subscribers on OpenAI’s new $500 a month Pro-500 plan. It was announced as “coming soon,” not live yet.
TL;DR
- Ultrafast mode targets around 300 tokens per second, about 8x the speed of normal ChatGPT generation, by running on Cerebras hardware instead of OpenAI’s standard GPU inference.
- Pricing runs 6x higher than standard inference, according to Sam Altman, making it a premium add-on rather than a replacement for everyday usage.
- Access requires the new Pro-500 tier, a $500 a month ChatGPT plan OpenAI introduced alongside Ultrafast at dev day, sitting above the existing $20, $100, and reactivated $200 plans.
- The $200 a month plan got less generous, with OpenAI cutting the usage included at that price point, effectively pushing power users toward the $500 tier to get comparable throughput.
- Ultrafast launched alongside other dev day features like the Dot assistant, GPT-6.1 Sol, Private Intelligence, and Codex in the cloud, all part of the same announcement wave.
- No live benchmarks or exact release date exist yet, since OpenAI described Ultrafast as forthcoming rather than shipped at announcement time.
Seven tools to build an app. Or just Remy.
Editor, preview, AI agents, deploy — all in one tab. Nothing to install.
Why is OpenAI charging so much more for speed?
Running inference on specialized hardware like Cerebras chips costs more than standard GPU-based serving, and OpenAI appears to be passing that cost directly to users rather than subsidizing it. Cerebras is known for wafer-scale chips built specifically to accelerate token generation, and pairing that hardware with OpenAI’s models is presumably what gets Ultrafast to roughly 300 tokens per second instead of the normal rate most ChatGPT users see.
The 6x price multiplier Altman mentioned suggests this isn’t meant to be a mass-market feature. It reads more like a tier for latency-sensitive workloads: voice agents, real-time coding assistants, or any application where waiting a few extra seconds per response actually breaks the user experience. For a developer or team where speed translates directly into product quality or revenue, paying a premium for faster tokens can pencil out. For a casual chat user, it almost certainly doesn’t.
How does the Pro-500 plan fit in?
OpenAI’s consumer and business pricing ladder now has four rungs: $20 a month (Plus), $100 a month (Pro), a reactivated $200 a month tier, and the new $500 a month Pro-500 tier. Ultrafast mode is gated entirely behind that top $500 plan, alongside access to OpenAI’s marketplace for third-party tools and apps.
What makes this notable is the change to the $200 plan. OpenAI reduced the amount of usage included at that price point, meaning people who want the level of access that $200 used to provide now effectively need to move up to $500. That’s a meaningful shift in how OpenAI is segmenting its subscriber base. Previously, $200 a month looked like the ceiling for individual power users. Now it’s a middle tier, and the real high-usage, high-speed experience sits at $500.
For context, OpenAI’s other major dev day announcement, the “Dot” always-on assistant, requires at minimum the $100 a month Pro plan. Ultrafast sits a full tier above that, making it one of the most expensive single features OpenAI has attached to a subscription plan to date.
Is Ultrafast mode worth the cost?
For most individual users, probably not. Paying 6x more for inference and committing to a $500 a month plan is a steep price for faster text generation when standard ChatGPT already returns answers in a few seconds. The use case is narrower: teams building latency-critical products, agentic workflows that chain together many model calls where delays compound, or real-time applications like voice assistants where every second of lag is noticeable to an end user.
It’s also worth noting that OpenAI has not published independent benchmarks or a firm release date for Ultrafast at the time of its announcement. The 300 tokens per second figure and the “8x faster” comparison come from OpenAI’s own dev day presentation, not from third-party testing. Until the feature actually ships and gets used at scale, it’s hard to know how consistent that speed is across different prompt types, model versions, or load conditions.
How does this compare to other fast-inference options?
Everyone else built a construction worker.
We built the contractor.
One file at a time.
UI, API, database, deploy.
Speed-focused inference isn’t new. Cerebras itself has offered fast inference services independent of OpenAI, and other providers have built businesses specifically around low-latency token generation using custom silicon. What’s different here is that OpenAI is bundling that kind of speed directly into its flagship consumer product rather than leaving it to third parties, and pricing it as a premium subscription feature rather than a pay-per-token API option alone.
That said, the transcript doesn’t indicate whether Ultrafast will also be available through OpenAI’s API for developers, or whether it’s strictly a ChatGPT subscription perk tied to Pro-500. If it stays subscription-only, that limits its usefulness for developers building products on top of OpenAI’s models, who typically care more about API-level pricing and throughput than consumer plan tiers.
What else was announced alongside Ultrafast?
Ultrafast was one of several announcements at the same dev day event. OpenAI also launched GPT-6.1 Sol, a cheaper model priced at $2 per million input tokens and $10 per million output tokens (compared to $10 and $50 for GPT-6 Astra), aimed at matching Astra’s coding performance at a lower cost. The event also introduced Dot, an always-on proactive assistant available to Pro and Business Premium users, Private Intelligence for confidential computing without OpenAI accessing underlying content, Codex running in the cloud for persistent coding sessions, a new Decisions API for structured decision-making tasks, and collaborative features like Spaces and Pages.
The breadth of announcements suggests OpenAI is pushing simultaneously on two fronts: cheaper, faster models for cost-sensitive use cases (GPT-6.1 Sol) and premium, high-performance tiers for users willing to pay more (Ultrafast, Pro-500). Ultrafast fits squarely into the second category.
Frequently Asked Questions
What chips does OpenAI’s Ultrafast mode run on?
Based on OpenAI’s dev day announcement, Ultrafast mode is built on Cerebras chips, which are purpose-built for fast AI inference rather than general GPU computing.
How much does Ultrafast mode cost compared to normal ChatGPT usage?
Sam Altman stated that Ultrafast mode costs 6 times more than standard inference, on top of requiring the $500 a month Pro-500 subscription plan.
Do I need the most expensive ChatGPT plan to use Ultrafast mode?
Yes. Ultrafast mode is restricted to the $500 a month Pro-500 tier, OpenAI’s highest subscription level, above the $20, $100, and $200 plans.
Is Ultrafast mode available right now?
At the time of its dev day announcement, OpenAI described Ultrafast as coming soon rather than immediately available, with no confirmed launch date.
Will Ultrafast mode be available through OpenAI’s API for developers?
The dev day announcement focused on ChatGPT subscription access via the Pro-500 plan. Whether a separate API option exists was not detailed in the announcement.

