Skip to main content
MindStudio
Pricing
BlogAbout
My Workspace
Claude Code system promptcontext engineering ClaudeAnthropic prompting guide

Anthropic Cut 80% of Claude Code's Prompt. Here's How to Trim Yours

Anthropic removed most of Claude Code's system prompt with no performance loss. Here's how to audit and trim your own CLAUDE.md and skills.

Edited by Luis Chavez-Mattos, Director of Product RSS
Anthropic Cut 80% of Claude Code's Prompt. Here's How to Trim Yours

What did Anthropic actually change in Claude Code’s system prompt?

Anthropic’s applied AI team rewrote the system prompt that ships with Claude Code, the instruction set that sits underneath every session alongside your own project context. According to an article published by one of Anthropic’s engineers on context engineering for Claude’s newer models, the team removed more than 80% of that prompt for models including Opus 4.5 and Sonnet 4.5, and reported no measurable drop in their evaluation scores. The rules that got cut were largely ones written to compensate for older, weaker models. As the models improved, those rules turned into restrictions that no longer matched what the model could actually handle on its own.

TL;DR

  • Anthropic deleted over 80% of Claude Code’s system prompt for newer models and saw no measurable performance loss on their evaluations.
  • Many old instructions existed to patch weaknesses in older models (like hard limits on code comments), and those same rules became unnecessary restrictions once the models got better.
  • Anthropic’s guidance is to use the smallest amount of context that still gets the result you want, not the largest amount you can fit.
  • A 2025 research paper found that longer inputs hurt performance on math, QA, and coding tasks across five models, even when the needed information was technically retrievable in context.
  • Loading MCP tool definitions only when needed cut token usage dramatically in Anthropic’s own testing and improved evaluation scores rather than hurting them.
  • Claude Code’s /doctor command audits your skills, memory files, and settings locally, flagging unused skills, duplicated instructions, and broken files you’d otherwise never notice.
  • The fix isn’t deleting everything. It’s keeping the context that defines intent and voice while cutting the step-by-step procedures the model no longer needs spelled out.
Cursor
ChatGPT
Figma
Linear
GitHub
Vercel
Supabase
goremy.ai

Seven tools to build an app. Or just Remy.

Editor, preview, AI agents, deploy — all in one tab. Nothing to install.

Why does adding more context sometimes make Claude worse?

The common assumption is that more information in context means better output, because the model has more to work with. Anthropic’s guidance pushes back on that directly: minimal context, not maximal context, is the goal. Minimal doesn’t mean short, either. If a task genuinely needs a detail, that detail stays. The problem is the accumulation of instructions that no longer serve a purpose but never get removed.

There’s research behind this beyond Anthropic’s internal numbers. A 2025 paper tested five models across math, question answering, and coding tasks, and found that performance dropped as input length increased, even in cases where the relevant information was still technically present and retrievable. The reported drops ranged from about 14% to as much as 85%, depending on the model and task. Those numbers came from controlled experiments on specific models, so they don’t translate directly into “trim your CLAUDE.md by X% and get Y% better results.” But they do establish that longer context carries a real cost, not just a theoretical one.

Anthropic’s own testing on tool definitions makes the point more concretely. When they tested loading MCP tool definitions only at the moment they’re needed, instead of all upfront, they reported an 85% reduction in token usage in their example. More notably, evaluation scores went up, not down. For Opus 4, MCP evaluation scores moved from 49% to 74%. For Opus 4.5, they moved from 79.5% to 88.1%. The model still had access to every tool. It just didn’t have to load all the instructions for tools it wasn’t using yet.

How do you know if your own Claude Code setup is bloated?

The same pattern that played out inside Anthropic’s system prompt tends to play out inside individual projects. Claude makes a mistake, you add a rule to your CLAUDE.md file or a skill to stop it from happening again. Months pass, you upgrade models, and the rule is still sitting there, sometimes duplicated in a slightly different form somewhere else. The model you’re running now may not need that guardrail at all, but nothing ever prompts you to go back and check.

Claude Code has a built in way to check: the /doctor command. Running it audits skills, memory files, plugins, settings, and installation health. It produces a table listing each skill, where it’s installed, how many times it’s actually been used, and an estimated token cost for just the metadata Claude reads to decide whether to load that skill at all (separate from the cost of loading the full instructions when it runs). In one documented run, this surfaced 18 older skills with zero recorded uses, while newer skills created days earlier were correctly left alone, since the audit accounts for age as well as usage.

One coffee. One working app.

You bring the idea. Remy manages the project.

WHILE YOU WERE AWAY
✓Designed the data model
✓Picked an auth scheme — sessions + RBAC
✓Wired up Stripe checkout
✓Deployed to production
Live at yourapp.msagent.ai

The audit also catches duplication. In the same example, a project map was found to be repeating information already covered by a separate routing map, with other sections duplicated for templates, references, and archiving. The audit quantified the overlap at roughly 2,250 characters, with an estimated savings of 563 tokens from cleaning it up. It also traces where memory is actually coming from: project-level CLAUDE.md, user-wide instructions, an auto-memory index, and a local machine-specific file, each loaded separately and each costing tokens even if they’re nearly empty. In one case, a leftover CLAUDE.md file in an unrelated desktop folder was still being loaded into a different tool, adding roughly 2,000 tokens the user didn’t know they were paying for. The same audit also caught two broken skill files, one with a bad filename, one with a formatting error in its metadata header that kept Claude from reading its description at all.

What should actually stay in a CLAUDE.md file or skill?

The useful filter is distinguishing between context the model cannot get anywhere else and procedure the model can now figure out on its own. Things like your voice, your goals, your audience, and how your specific business or project works belong in context, since Claude has no way to infer or look those up. The intent behind a request matters too. Knowing a document is meant for students, for example, helps the model choose simpler language and better analogies, and that’s worth keeping explicit.

What tends to be safe to cut is granular procedure: rigid formatting rules, repeated instructions that exist in three places instead of one, and step by step processes that only apply to specific situations (a detailed video editing procedure doesn’t need to load when you’re asking for help with a script). The comparison that captures this well: giving an experienced model a rigid checklist is like handing a seasoned designer the instructions meant for someone making their first presentation. The designer still needs to know the audience and the goal. Dictating the placement of every element before they’ve seen the material just prevents them from contributing something better than what you specified.

In practice, this means testing a trimmed version against your current setup. One useful exercise is comparing output from a fully loaded setup against a stripped down one on the same task, then merging the best of each (the branding and structure you want to keep, applied as a lighter skill, paired with the more flexible content organization a less constrained run produced).

Frequently Asked Questions

What is Claude Code’s system prompt?

It’s the base set of instructions that ships with Claude Code and governs how the model behaves by default, sitting alongside whatever project-specific context, CLAUDE.md files, and skills a user adds on top.

Did cutting the system prompt actually hurt performance?

Anthropic reported no measurable loss on their evaluations after removing over 80% of the prompt for newer models like Opus 4.5 and Sonnet 4.5. Many of the removed rules existed to compensate for limitations in older models that no longer apply.

What does /doctor do in Claude Code?

It audits your local setup: skills, memory files (project, user, and local level), plugins, settings, and installation health. It reports usage counts per skill, flags duplicated instructions across files, estimates token costs, and identifies broken or misconfigured files.

Does “minimal context” mean writing shorter prompts?

Not necessarily. Anthropic’s guidance is to use the smallest amount of context that still produces the result you want, which can still be long if the task genuinely requires that detail. The goal is removing instructions the model no longer needs, not shortening instructions it does.

How often should I clean up my skills and CLAUDE.md files?

Plans first. Then code.

PROJECTYOUR APP
SCREENS12
DB TABLES6
BUILT BYREMY
1280 px · TYP.
yourapp.msagent.ai
A · UI · FRONT END

Remy writes the spec, manages the build, and ships the app.

There’s no fixed rule, but the practice of periodically auditing accumulated instructions, rather than only ever adding to them, is the core idea behind this shift in guidance. Running /doctor periodically and checking for unused skills or duplicated rules is a practical way to do that.

Editorial standards

Presented by MindStudio

No spam. Unsubscribe anytime.