Skip to main content
MindStudio
Pricing
BlogAbout
My Workspace
Claude Opus 5.5Opus 5.5 reviewClaude token efficiency

Claude Opus 5.5: Why It Needs Fewer Tokens to Finish Work

Claude Opus 5.5's token efficiency, pricing, and writing steerability explained through a real 89-million-token build and cost breakdown.

Edited by Luis Chavez-Mattos, Director of Product RSS
Claude Opus 5.5: Why It Needs Fewer Tokens to Finish Work

What makes Claude Opus 5.5 more token-efficient than previous Opus models?

Claude Opus 5.5 gets more done per token because it needs fewer passes, retries, and clarifying exchanges to produce a finished result. Anthropic cut list prices by 20% compared to Opus 5 ($4 per million input tokens, $20 per million output tokens), but the bigger story is behavioral: the model reportedly requires less back-and-forth to reach a correct answer. Anthropic says typical workloads cost about 40% less on Opus 5.5, a number that combines the price cut with the model simply chewing through fewer tokens to finish the same task.

TL;DR

  • Opus 5.5 charges $4/$20 per million input/output tokens, a 20% price cut from Opus 5 and roughly 60% cheaper than GPT-5.1’s $10/$50 rates.
  • Anthropic claims typical workloads run about 40% cheaper on Opus 5.5, driven by a mix of lower prices and the model needing fewer tokens to finish the same work.
  • A real-world test case: a complex Lego-model build (logo turned into a 514-piece set with instructions, parts list, and geometry checks) used 89 million tokens, costing an estimated $44 at API rates but consuming only about 1% of a $50/week subscription allowance.
  • Token count alone doesn’t tell the full efficiency story: input, output, cache, and context tokens all carry different rates, so judging “efficiency” requires looking at the whole job, not a single screenshot or clip.
  • Writing steerability is back: after community frustration with Opus 5 and GPT-5.1 era models fighting back on edits, Anthropic specifically targeted clearer writing and better instruction-following as a fix in 5.5.
  • Coding-for-visuals is a growing pattern: Anthropic models increasingly use tools like Three.js to generate and revise 3D scenes through code, which makes targeted edits (change the hat, keep the glasses) far cheaper than regenerating a whole scene.
  • Long-running autonomous jobs do better with explicit stop conditions: without a clear “definition of done,” Opus 5.5’s persistence can turn into token-hungry overreach.

Remy doesn't write the code. It manages the agents who do.

R
Remy
Product Manager Agent
Leading
Design
Engineer
QA
Deploy

Remy runs the project. The specialists do the work. You work with the PM, not the implementers.

How much does a real Opus 5.5 task actually cost?

The clearest illustration from hands-on testing is a single, complex creative task: converting a logo (a figure in a beanie and glasses) into a buildable Lego model. The output wasn’t just an image. It included an animated build sequence, a 63-page instruction booklet with 58 build steps, a full parts list, a BrickLink “wanted list” for ordering real pieces, a geometry-check file verifying brick connections, and an LDR file confirming the model could physically be assembled. The finished design matched the original logo and came out to 514 pieces.

That entire job consumed 89 million tokens. At API pricing, that would run at least $44. On a flat-rate subscription plan (around $50/week in this case), the same task consumed only about 1% of the weekly usage allowance, meaning the effective cost to the user was closer to 50 cents. Scaled to a $200/month plan, 89 million tokens would represent roughly a quarter of the monthly price if paid at raw API rates, yet it barely dented a weekly usage cap.

This gap between “what it would cost on the API” and “what it actually consumes against a subscription” is the practical argument for why task efficiency matters more than headline token prices.

Why does token efficiency matter more than per-token pricing?

Per-token price is only half the equation. The other half is how many tokens a model burns to get to an acceptable result. A model that’s 20% cheaper per token but needs twice as many retries to satisfy a request isn’t actually cheaper to use. Anthropic’s own framing (a roughly 40% drop in typical workload cost) explicitly separates the price reduction from the efficiency gain, because the two don’t automatically move together.

Several companies referenced in Anthropic’s own release materials describe this pattern directly: GitHub and Lovable reported Opus 5.5 needing fewer steps to complete tasks, and Spotify reported completing comparable work more cheaply and faster. That lines up with anecdotal reports circulating elsewhere from individual users noticing the same kind of drop in retries and token churn.

The practical lesson for anyone evaluating a model: don’t just compare sticker prices per million tokens. Look at the full token bill for a complete, real task, across input, output, cached, and context tokens, because those are billed at different rates and a model’s actual efficiency only shows up when you add up the whole job.

Is Opus 5.5’s writing actually easier to work with?

For a stretch of releases before 5.5, a meaningful chunk of user complaints centered on Claude’s writing behavior becoming harder to steer. Writer Bram Cohen published a widely shared piece criticizing Claude’s argumentative tendencies, and Anthropic’s own GitHub issue tracker contains user reports describing writing instructions being acknowledged by the model and then ignored anyway. That’s a real cost: a model that quietly overrides explicit instructions forces the user to re-explain, re-check, and re-edit, which burns both time and tokens.

REMY IS NOT
  • ✕a coding agent
  • ✕no-code
  • ✕vibe coding
  • ✕a faster Cursor
IT IS
✓a general contractor for software

The one that tells the coding agents what to build.

Anthropic’s 5.5 announcement names this directly, citing writing clarity and instruction adherence as one of the top fixes targeted in the release. In hands-on use, the difference shows up as steerability: asking the model to keep a specific tone, preserve an intentional note of uncertainty, or simplify a sentence without stripping its substance, and having that correction land on the first try rather than requiring several rounds of pushback.

This matters because AI writing edits aren’t neutral. Making a paragraph “simpler” can quietly delete the nuance that made an argument worth reading. Making a tone “warmer” can soften a decision that needed to sound firm. Making a sentence “more confident” can erase uncertainty the writer deliberately included. Good revision preserves intent; sloppy revision just produces something that reads smoothly while losing the point. Steerability is what determines which one you get.

How does Opus 5.5 handle long, unattended tasks?

Anthropic’s release materials cite an example from Cleo, where an engineer left Opus 5.5 running an unattended task across six code repositories for 18 hours. Multi-hour autonomous runs aren’t new to this model generation, but they raise a specific efficiency question: a model left to work independently for that long is also making a long string of unsupervised decisions, any of which can go sideways.

Hands-on testing with a few overnight jobs surfaced a practical pattern: Opus 5.5 performs better on long-running tasks when given explicit stop conditions and a clear definition of “done” up front. Left open-ended, the model’s tendency toward persistence can turn into unnecessary token consumption as it keeps refining or second-guessing a result that was already good enough. Specifying constraints in advance (for the Lego build, that meant defining piece count and complexity before the model started) prevents the model from having to model the entire problem space itself and then decide independently when to stop pushing.

How can you tell if a model is actually more token-efficient for your own work?

Measure the whole job, not a single output. A good-looking screenshot, a clean paragraph, or one successful code change doesn’t tell you whether a model is efficient, because the real cost includes every retry, every re-explanation, and every token spent on context and caching along the way. To evaluate efficiency for a specific use case: run a complete, representative task end to end, record the full token bill across input, output, and cache categories, compare that to what the same task would have cost in a previous model version, and check how much of a subscription’s usage allowance it consumed in practice versus raw API pricing. Because models are “spiky,” meaning they perform unevenly across different kinds of tasks, this kind of individual measurement matters more than any single published benchmark.

Frequently Asked Questions

How much does Claude Opus 5.5 cost per token?

Opus 5.5 is priced at $4 per million input tokens and $20 per million output tokens, a 20% reduction from Opus 5 and about 60% below GPT-5.1’s $10/$50 rates.

What does Anthropic claim about Opus 5.5’s cost savings?

Anthropic states that typical workloads cost roughly 40% less on Opus 5.5, attributing the savings to a combination of lower per-token pricing and the model requiring fewer tokens overall to complete tasks.

What was the actual token cost of the Lego build example?

The complex Lego-model generation task, which produced a build animation, instructions, parts list, and geometry-check files, used 89 million tokens, equivalent to roughly $44 at API pricing but only about 1% of a weekly subscription usage allowance.

Did Anthropic fix the writing complaints people had about earlier Claude models?

Anthropic’s Opus 5.5 announcement specifically cites clearer writing and better adherence to writing instructions as a targeted fix, responding to public feedback about earlier models being harder to steer and prone to ignoring explicit instructions.

Why does Opus 5.5 use coding tools for visual tasks?

Anthropic models increasingly generate and revise visual scenes by writing code (for example, using Three.js) rather than regenerating images from scratch, which lets the model make precise, targeted edits to a scene while leaving everything else intact, saving tokens compared to full regeneration.

Editorial standards

Presented by MindStudio

No spam. Unsubscribe anytime.