Don’t raise Opus 5.5 effort yet. Audit the prompts first.
Lance Martin posted the tip on the same day Anthropic shipped Claude Opus 5.5. Two thousand likes in a few hours. The command is one line: /claude-api prompt-audit
Most people skipped it and went straight to /effort xhigh. That is the wrong order.
Opus 5.5 already thinks before every reply. Thinking cannot be turned off on this model. Effort only turns the volume up. If your CLAUDE.md still says “think step by step,” “verify twice,” or “MUST NEVER skip the six-step checklist,” you are paying for a ritual the model already performs. Anthropic measured that tax. On a 44-ticket support benchmark, moving from Opus 4.8 to Opus 5.5 at low effort cut cost about 18 percent. Running the prompt audit cut another 9 percent. Same tasks. Fewer tool calls. The model stopped repeating itself.
This post is the missing middle. Not another “Opus 5.5 is out” announcement. The official pages already own that query. This is what to do to your repo tonight so the new default model does not drown in instructions written for an older one. The source procedure lives in Anthropic’s public skills repo at skills/claude-api/shared/prompt-audit.md.
Why Opus 5.5 still overthinks on day one

The model drop landed on 22 September 2026. Official claim set is simple enough: Fable 5.1 level work on most tasks, about 40 percent cheaper to run than Opus 5, more than 30 percent faster output. Default effort moved from high (Opus 5) to medium (Opus 5.5). That last line is the one people missed.
Effort names do not travel across models. Medium on 5.5 is not medium on 5. Anthropic’s own coding evals show 5.5 at medium matching or beating 5 at high. On FrontierCode, 5.5 peaks around medium and high actually slips. A thread on r/ClaudeAI the same night put Terminal-Bench, FrontierCode, and CursorBench next to dollar costs. Xhigh added 89 percent cost on Terminal-Bench for 2.2 accuracy points. On FrontierCode it cost more and scored worse.
So the overthinking complaint is real. It is also often self-inflicted. You carried /effort high or /effort xhigh from the last model. You left “be thorough” and “double-check every file” in CLAUDE.md. The new model treats those lines as extra work, not flavor text.
Swyx ran Latent Space’s AINews draft through 5.5 next to GPT-6 Sol and called 5.5 the new default because it was more concise and tasteful with less slop. Simon Willison said the same after a side-by-side. Those are writing tests. Coding sessions behave differently. Coding sessions load your skills, your agent.md, your CLAUDE.md, and every MCP description you left connected. That pile is what the audit is for.
If you want the mental model for how Claude Code even loads this stuff, the walkthrough of the learn-claude-code repo is still the cleanest map of the loop, the tools, and on-demand skill loading: How Claude Code works — the GitHub repo that reverse-engineers it step by step.
What the Opus 5.5 prompt audit actually inspects
I opened the official skill, not a blog summary of it. The file lives in Anthropic’s public skills repo at skills/claude-api/shared/prompt-audit.md. The first paragraph is a warning: if you arrived via /claude-api prompt-audit, execute the steps. Do not summarize them back to the user. Official prompting guidance for this model is also on Anthropic’s docs for Prompting Claude Opus 5.5.
Two artifacts. Always both.
1. An audit report. Every finding gets a file:line location, the named pattern it matches, why that pattern is obsolete on the target model, and a confidence level.
2. A proposed diff. Concrete edits. The skill is explicit: propose, never apply without consent.
The prime directive inside the file is the line most listicles will skip. Distinguish cruft from load-bearing content. A finding you cannot tie to a named pattern, with a reason grounded in the target model’s documented behavior, is not a finding. “Make it shorter” is not the job. “Every token earns its place” is.
The scanner walks four groups.
Dated prompt text. Caps-lock pressure (MUST, NEVER, CRITICAL). Hedges that now read as permission to under-deliver (try to, if possible). Trait claims (don't be too verbose). The old scaffolds: “think step by step,” <scratchpad> tags, “use the think tool to plan,” “show your thinking,” forced JSON prefill, “summarize progress every N tool calls,” budget_tokens, non-default temperature, tool_choice forced to a named tool. Step-by-step choreography for judgment tasks. Prohibition lists with no business reason. Identity stubs. “Reminder:” lines re-inserted every few turns.
Brittle skill files. A SKILL.md that explains what the model already knows. Rules copied from one bad session. Hardcoded paths and version numbers. Trigger lists that grow every week.
Tool descriptions. One-liners with no when-not-to-use. CRITICAL: You MUST use this tool when... Overlapping tools. Behavior smuggled into a description (“after showing results, always recommend…”).
Request config. Deprecated API fields. Cache-hostile ordering (timestamps and UUIDs above stable content). Specialist sub-agents that duplicate each other.
If that list feels familiar, it should. A lot of public Claude Code setups from early 2026 are a museum of those patterns. The 65-line CLAUDE.md pattern existed because agents over-edited. The fix then was tighter rules. The fix now is fewer rules, written as constraints with reasons, not as a second personality: 91k stars — the CLAUDE.md file that fixes AI coding.
The numbers that should change what you type tonight
Treat these as published measurements, not a promise about your repo.
On Anthropic’s 44-ticket support benchmark, the prompt had a mandatory six-step procedure, a scratchpad rule, a verify-twice rule, and contradictory instructions. Opus 5.5 at low effort already beat the Opus 4.8 baseline by about 18 percent on cost. The audit removed those rituals and cut another 9 percent. Combined: about 25 percent under the old starting point. Anthropic says treat it as an example. Fair. The mechanism is still the part that transfers. Rituals create extra tool calls. Extra tool calls create extra output. Output is $20 per million tokens on 5.5. Input is $4. Cache reads are $0.20. You want fewer turns and a high cache share, not a longer sermon in CLAUDE.md.
A separate earlier test on Opus 5 (same skill family) planted one anti-pattern at a time across six legacy prompts. After prompt-audit, cost fell 14.6 percent and accuracy rose 5.3 percent on average. “Verify twice” duplicated order lookups. “Be maximally thorough” became dozens of knowledge-base searches the ticket did not need.
Price list for the new model, so the effort debate has units. The cost post lists Opus 5.5 at $4 per million input and $20 per million output. Cache read is 5 percent of input. A 120K-context, 40-turn task with 90 percent cache is $1.62 in their worked example. Same task in 25 turns is $1.02. Sixty thousand output tokens alone are $1.20. High effort that adds 20K thinking tokens costs about $0.40. That $0.40 is cheap if it saves a retry. It is waste if medium would have finished.
Official recommendation that follows from those numbers: start 5.5 at medium for scoped daily work. Move to high when medium stalls. Use low for mechanical edits. Reserve xhigh and max for work where you measured a gain. Changing effort mid-session clears the prompt cache, because effort sits in the prefix the cache matches. Do the audit first. Then pick one effort and leave it alone for the task.
The 8-minute test you can run on your own repo
You need Claude Code with the claude-api skill available. Lance updated that skill the day 5.5 shipped. If /claude-api is missing, add Anthropic’s official skills marketplace and install it. Then stay in the repo you actually work in, not a toy folder.
Minute 0 to 1. Check the model and the dial.
/model/effort statusIf the model is still Opus 5, switch. If effort is still high or xhigh from last month, set medium before you judge anything.
/model claude-opus-5-5Minute 1 to 2. Snapshot cost on a short, familiar task you already trust. One bug fix. One test run. End with /usage or /cost. Write down turns, output tokens, cache share. That is your before.
Minute 2 to 6. Run the audit.
/claude-api prompt-auditDo not accept the diff blind. Read the report. Keep constraints only you know: repo layout, deploy rules, “never force-push main,” the house style that is not a vibe, it is a contract. Delete the rest of the museum. Pressure language. Scratchpad rules. Verify-twice. Think-step-by-step. Six-step dances for judgment work. Identity stubs. Reminder loops.
Minute 6 to 8. Re-run the same short task. /usage again. You are looking for fewer turns and fewer output tokens at the same quality, not a miracle. If quality drops, put back one load-bearing rule at a time. The skill itself says to verify before/after and re-add minimally if something regresses.
A clean CLAUDE.md after this pass usually looks boring. Who you are. How the repo is laid out. What “done” means. When to stop and ask. What is destructive. Under 200 lines, which is also Anthropic’s own advice in the cost post. Karpathy’s four rules still belong here if they are written as constraints, not as a second system prompt. The OfficeCLI skill pattern is a useful contrast: that SKILL.md is long because it teaches a tool the model does not already know. Your house style file is the opposite. The model already knows how to write TypeScript. It does not know that your payments package cannot import the CLI package: OfficeCLI — AI agent office automation without installing Office.
While the model is mid-task you can still type. “Also keep the old endpoint names as aliases.” You do not need to restart the run. You do need to stop asking it to think. Thinking is already on. If you want less thinking, lower effort. Prompt lines lose that fight.
Effort vs cleanup: pick the cheaper lever
People reach for xhigh because it feels like buying intelligence. On 5.5 that purchase is uneven.
The r/ClaudeAI tables from launch night are the honest ones. Terminal-Bench: medium 57.6 percent at $2.94, high 64.2 percent at $3.88, xhigh 66.4 percent at $7.35, max 64.8 percent at $11.24. High is the last step that still pays. Xhigh almost doubles the bill for a sliver. Max spends more than xhigh and can score worse. FrontierCode: medium 54.6 percent at $0.80 is the peak. High, xhigh, and max do not help. CursorBench: high helps a bit, xhigh is a flat line at nearly double the cost.
So the rule that survives contact with those tables is narrow. Stay on medium. If the session stalls on a real multi-file problem, go to high. Do not live on xhigh because a launch thread said extra high is for agents. Extra high is for the slice of agent work where you already measured a gain.
Cleanup is cheaper than that whole ladder. Removing a verify-twice rule costs nothing at runtime and can delete a whole extra tool-call loop. Removing four overlapping MCP servers costs nothing at runtime and shrinks the prefix you pay to cache. Unneeded MCP connections are on Anthropic’s own waste list, next to mid-session model switches and compacting while you are still debugging.
One more migration footgun. Thinking cannot be disabled on Opus 5.5 at any effort. Requests that send thinking: {"type": "disabled"} return 400. Forced tool_choice on the new family can 400 as well. If you wrap Claude Code in your own harness, run the same audit against the request-building code, not only against markdown. The skill is built for that path too.
What this means if you build agents for a living
The last year trained a bad habit. When the model got worse, we added a rule. When it got verbose, we added “don’t be verbose.” When it skipped tests, we added a six-step definition of done. The file grew. The model got stronger. The file did not shrink.
Opus 5.5 is the first Claude in this line that makes that habit expensive in public. Official writing guidance says delete “think carefully.” Official cost guidance says run prompt-audit on skills and CLAUDE.md. Official effort guidance says the default is medium and the names do not port. Three documents, one message. The model improved. Your prompts did not.
Keep the parts only you can write. The architecture of the repo. The incidents that actually happened. The APIs that are allowed to talk to each other. The sentence that defines done for this team. Give large work to subagents and demand evidence back, which is the other thing early testers kept repeating. Then get out of the way.
If you try only one thing from this page, make it the command, not the effort slider.
/claude-api prompt-auditRead the diff. Keep what you can defend. Measure one task with /usage. Raise effort only after the file is clean.
Common questions about Opus 5.5 prompt audit
Should I run an Opus 5.5 prompt audit before changing effort?
Yes. Effort multiplies whatever the prompt already asks for. A dirty CLAUDE.md at high effort is a more expensive dirty CLAUDE.md. Audit, measure one familiar task at medium, then raise effort only if that task stalls.
What does /claude-api prompt-audit change in Claude Code?
It scans the prompt surface in the working directory: CLAUDE.md, skills, agent files, tool descriptions, and application code that calls the Claude API. It returns a report plus a proposed diff. It does not silently rewrite your repo.
Why does Opus 5.5 ignore “think step by step”?
It does not ignore it. It obeys it, on top of thinking that is already on at every effort level. That is why the line is in the deletion list. You wanted a nudge. You got a second thinking pass.
Which Opus 5.5 effort setting should I use in Claude Code?
Medium for scoped daily work. High when medium misses a layer. Low for mechanical edits. Xhigh and max only where your own evals show a gain. Changing the setting mid-task clears the prompt cache.
Will deleting rules make the agent sloppy again?
It can, if you delete a constraint only you know. The audit’s own rule is the guardrail. If you cannot name the pattern and the reason it is obsolete on 5.5, leave the line. Put back anything that regresses. The goal is a short file of load-bearing facts, not a blank file.
1 comment