GPT-6.1 Soul Pricing: How It Undercuts GPT-6 Astra by 80%
GPT-6.1 Soul costs a fifth of GPT-6 Astra with near-Astra performance. Full per-token pricing and caching discount breakdown inside.

What does GPT-6.1 Soul cost compared to GPT-6 Astra?
GPT-6.1 Soul runs at $2 per million input tokens and $10 per million output tokens, with cached input priced at just 10 cents per million tokens. GPT-6 Astra, OpenAI’s flagship model, costs $10 per million input tokens and $50 per million output tokens. That makes Soul roughly a fifth of Astra’s price on standard input and output, an 80% saving, while OpenAI positions it as close to Astra in actual performance.
TL;DR
- GPT-6.1 Soul prices in at $2/$10 per million input/output tokens, against $10/$50 for GPT-6 Astra, an 80% discount for a model OpenAI claims performs close to the flagship tier.
- Cached input tokens for Soul cost only 10 cents per million, a 95% discount off the standard input price, which matters most for agents that repeatedly reread system prompts and conversation history.
- Soul plays the same role Opus-tier models have played elsewhere: a smaller, cheaper sibling to the top model (Astra) that closes much of the performance gap while cutting cost dramatically, similar to how Opus 5.5 narrowed the distance to Fable.
- Agent workloads will likely run mostly on cached tokens, since system prompts and context get reread constantly. One builder reported that out of 250 million tokens processed, about 96% were already cached, which collapses real-world costs far below the headline per-token rate.
- An “ultrafast” tier of Astra also launched, running at roughly 8x the speed for 6x the price, aimed at latency-sensitive use cases rather than cost-sensitive ones. OpenAI said an ultrafast version of Soul is coming.
- For most agent builders on OpenAI’s stack, Soul 6.1 is shaping up to be the new default model, trading a small amount of peak capability for a large reduction in running cost.
- ✕a coding agent
- ✕no-code
- ✕vibe coding
- ✕a faster Cursor
The one that tells the coding agents what to build.
Why does the caching discount matter more than the sticker price?
The headline numbers ($2 in, $10 out) already make Soul attractive next to Astra. But the real story for anyone running agents is the caching price: 10 cents per million tokens for cached input, a 95% discount off standard input pricing.
Agent systems don’t process fresh context on every call. They reread the same system prompt, the same tool definitions, and the same accumulating conversation history over and over as a task runs. That repeated context is exactly what prompt caching is built to discount. If an agent loop sends the same 10,000-token system prompt on every step of a long-running task, only the new tokens at the end of that prompt need to be billed at full price. Everything that matches a previous call gets billed at the cached rate.
In practice, this means the effective cost of running an agent on Soul can be far lower than multiplying token counts by the standard rate. A workload that looks like 250 million tokens on paper can end up mostly billed at the cached rate if 90%+ of those tokens are repeated context rather than new generation. That’s the kind of ratio builders are already seeing with other frontier models in this same generation, and it’s reasonable to expect similar cache hit rates on Soul for typical agent loops.
How does Soul’s performance compare to Astra?
OpenAI is positioning GPT-6.1 Soul as its Opus-equivalent: a mid-tier model sized and priced below the flagship, but close enough in capability that most workloads won’t need to pay flagship prices. The comparison OpenAI invited was to Opus 5.5, which reportedly closed much of the gap to its own flagship sibling, Fable.
The caveat is that this closeness likely breaks down at the extremes. When a task pushes maximum reasoning depth, maximum thinking budget, or otherwise stresses a model’s ceiling, the top-tier model (Astra) is still expected to edge out Soul. For the bulk of everyday agent and coding tasks, though, the performance difference is apparently small enough that the 80% price cut looks like a clear win.
Is GPT-6.1 Soul worth switching to for agent workloads?
For anyone running agents on OpenAI’s models, the pricing structure makes Soul 6.1 the practical default rather than a budget fallback. A few reasons line up in favor of that:
- The input/output pricing alone cuts cost by roughly 80% versus Astra for comparable work.
- The caching discount (95% off input) compounds that savings further for any workload with repeated context, which describes most agent harnesses, coding assistants, and multi-step task runners.
- The performance gap to Astra appears narrow outside of maximum-effort reasoning tasks.
Where it’s not worth switching: tasks that specifically need the ceiling Astra provides, deep multi-step reasoning chains where every point of model capability matters, or workloads where latency rather than cost is the binding constraint. For those cases, OpenAI also introduced an “ultrafast” tier, currently available on Astra and coming to Soul, running at around 8x normal speed for 6x the price. That tier is aimed at situations where a user is actively waiting on output, like live editing or interactive demos, not at reducing cost.
How does this fit into OpenAI’s broader pricing strategy?
Soul’s pricing lands alongside a broader DevDay pattern: OpenAI tiering its lineup more aggressively by both capability and price. Astra sits at the top as the expensive, highest-capability model. Soul sits underneath as the cost-efficient workhorse. Ultrafast variants sit alongside both for latency-sensitive cases, charging a speed premium rather than a capability premium.
That same event also saw OpenAI adjust its subscription tiers, introducing a $500/month Pro plan with 25x the usage of Plus and access to ultrafast inference, while the existing $200/month Pro plan’s usage multiplier was reduced from 20x Plus down to 10x Plus. Read together, the message is consistent: OpenAI is pushing high-volume users, whether on the API or on subscriptions, toward paying for exactly the tier of speed and capability they need rather than defaulting everyone to the flagship.
For API users building agents, the practical takeaway is that Soul 6.1’s combination of a lower base rate and a steep caching discount makes it the model worth defaulting to, with Astra reserved for the specific steps in a pipeline that actually need the extra ceiling.
Frequently Asked Questions
How much cheaper is GPT-6.1 Soul than GPT-6 Astra?
Soul costs $2 per million input tokens and $10 per million output tokens, versus $10 and $50 for Astra. That’s roughly an 80% reduction in standard token pricing.
What is the cached input price for GPT-6.1 Soul?
Cached input tokens cost 10 cents per million, a 95% discount off Soul’s standard input rate. This applies to context that repeats across calls, such as system prompts and conversation history in agent loops.
Is GPT-6.1 Soul as capable as GPT-6 Astra?
OpenAI claims Soul performs close to Astra on most tasks, similar to how Opus 5.5 narrowed the gap to its flagship sibling. The expected shortfall shows up mainly on tasks that push maximum reasoning depth or thinking budget, where Astra still has an edge.
What is the “ultrafast” tier and does it apply to Soul?
Ultrafast is a faster inference option, roughly 8x the speed for 6x the price, currently available for Astra. OpenAI said an ultrafast version of Soul is coming, which would let latency-sensitive workloads get faster output from the cheaper model too.
Should agent builders default to Soul instead of Astra?
For most agent and coding workloads, yes. The lower base price plus the heavy caching discount make Soul the more cost-effective choice for repeated, high-volume use. Astra remains the better option for tasks that specifically require maximum model capability.


