Why Claude Code costs more the moment you start measuring it

2026-09-26

In 2026 the Claude Code MCP token cost worth worrying about is rarely the idle tool list. Tool definitions are deferred by default, and Opus 5.5 cache reads cost $0.20 per million tokens. The money goes on cache misses, and on one quiet trap: any local proxy that sets ANTHROPIC_BASE_URL switches tool search off. ENABLE_TOOL_SEARCH=true is the fix.

cost-xray showed up on Trendshift this week with 49 live mentions, from the same Tigerless Labs team behind AutoHarness. The pitch is great: see exactly what Claude Code sends to the API, and what each piece of it costs. The README promises it “doesn’t change what your agent does, its results, or its cost.”

I read the shell wrapper it installs. The whole trick is one line: ANTHROPIC_BASE_URL="http://127.0.0.1:$p" command claude. Then I read the Claude Code MCP docs, which say, word for word: “Configurations without tool search include a custom ANTHROPIC_BASE_URL.”

So the tool you install to find MCP waste can switch on the MCP waste it reports. Not a scandal. It is a one-line fix, and the same bug hit two other proxy projects this month. But if you read a cost-xray dashboard without knowing this, you will “fix” a problem that only exists while you are watching.

What does Claude Code actually send to the API?

Mind map of Claude Code MCP token cost: cache misses, ANTHROPIC_BASE_URL tool-search trap, ENABLE_TOOL_SEARCH fix

Start with one plain fact. The model has no memory between messages. Every time you press enter, Claude Code sends the whole thing again: its instructions, the list of tools it may use, your CLAUDE.md, and every message and file read so far.

That pile is measured in tokens (chunks of text, roughly three quarters of a word each, the unit you are billed in). The most the model can read at once is the context window (think of it as the desk it works on; anything that does not fit falls off).

Resending everything every turn sounds ruinous. It would be, without prompt caching (the API keeps a copy of the start of your request for a few minutes, and if your next request starts the same way, re-reading that part is billed at a steep discount). The shared start is called the prefix (the part of the request that did not change since last time).

Picture a print shop. New pages cost full price. Pages that are already sitting on the glass from your last job cost almost nothing. But if you change one early page, every page after it counts as new. That is the rule the whole bill hangs on: the match is exact, and a change near the top reprices everything below it.

Claude Code orders the request to protect that. The system prompt and tool definitions go first, because they rarely change. Project context like CLAUDE.md goes next. The conversation goes last, since it grows every turn. The official prompt-caching page spells it out: a change to the conversation keeps the rest cached, while a change to the system prompt layer “invalidates everything.”

What is the Claude Code MCP token cost per turn in 2026?

MCP (Model Context Protocol, the standard plug that lets Claude Code talk to outside tools like GitHub, a browser, or Notion) adds tools. Each tool comes with a schema (a small block of JSON describing its name, what it does, and what inputs it takes). The old worry was simple: every connected server stuffs all its schemas into every request, used or not. Guides from earlier this year measured Playwright alone at about 3,400 tokens.

Two things changed since those guides were written.

First, tool search (Claude Code’s habit of sending only tool names up front and loading a full schema only when Claude reaches for that tool) is on by default on supported models from v2.1.232. With it on, an idle server costs a line of names, not a wall of JSON.

Second, the price of re-reading fell. Opus 5.5 lists at $4 per million input tokens and $20 per million output. Cache reads dropped from $0.50 to $0.20 per million, which Anthropic’s own pricing table calls 0.05x the base input price. Writing to the cache costs $5 per million with a five-minute lifetime, or $8 with a one-hour lifetime.

Here is what that does to the Claude Code MCP token cost of a single turn, input side only, at Opus 5.5 list prices:

Prefix sizeWarm turn (cache read)Cold turn, 5-min writeCold turn, 1-hour write
20,000 tokens$0.004$0.10$0.16
60,000 tokens$0.012$0.30$0.48
150,000 tokens$0.030$0.75$1.20

A cold turn (the cache expired or got broken, so the whole prefix is written again) costs 25 times a warm one on the five-minute rate and 40 times on the one-hour rate. Nothing else in the bill swings that hard.

Now run a day. Say 150 turns on a 150,000-token session. All warm: about $4.50 of input. Six misses: $8.82. Twenty misses: $18.90. Same code, same prompts, four times the bill. Misses decide it.

And misses are easy to cause. The Claude Code docs list the usual suspects: switching models, turning on fast mode, compacting, upgrading, and connecting or removing an MCP server when its tools are loaded into the prefix. There is also the plain coffee break. On a Claude subscription the main conversation gets a one-hour cache. On an API key, or once you are paying with usage credits, it defaults to five minutes. Walk away for six minutes on an API key and your next message rewrites everything.

What cost-xray shows that /usage does not

Claude Code already tells you a lot for free. Run /usage on v2.1.251 or later and you get a “Prompt cache (main)” line with your hit rate, your misses, and whether the cache is warm right now. From v2.1.260 it even guesses the cause of the last miss, for example “tool definitions changed.” Run /context to see what fills the window. Check those first.

What /usage will not do is put a price on each piece. cost-xray does. It runs mitmproxy (an open-source tool that sits between a program and the internet and records the traffic passing through) on your own machine. Claude Code talks to it on 127.0.0.1, it forwards to Anthropic, and it saves a redacted copy with keys and cookies stripped before anything reaches disk.

Then a separate step prices the capture. Every token gets sorted into fresh input, cache read, cache write, or output, and pinned to the thing that caused it: the system prompt, one tool’s schema, one MCP server, a single Read call and its result. Log readers like ccusage see the transcript after the fact, and the transcript never contains the system prompt or the schemas. The README estimates that hidden prefix can be half the context or more. Only the wire has it.

It also stores the traffic cleverly. A long session resends its whole history every turn, so cost-xray keeps each unique block once and records small per-turn changes. It flags “unused MCP waste” too, meaning servers whose schemas ride along on every request but never get called.

I cloned it at commit 5d69deb (1 September 2026) and ran its test suite. 223 tests passed. Then I built a request shaped like a Claude Code turn, with a system prompt, 12 built-in tools and 61 MCP tools across three servers, and fed it through cost-xray’s own analyze and reconcile_turn functions. MCP schemas came to 24,360 of 38,479 tokens. Only the GitHub server got called. Playwright and Notion sat unused at 13,986 tokens between them. The input side of that turn cost $0.0088 warm and $0.3066 cold on a one-hour write.

To be plain about it: those schemas are synthetic and I set the usage numbers by hand. This environment cannot run your login, so this is the pricing engine on a realistic shape, not a live capture of your session.

Yes, according to Anthropic’s own docs, and this is the part that matters.

A reverse proxy (a local stand-in address that receives your requests and passes them on to the real server) is the easy way to watch Claude Code. You point ANTHROPIC_BASE_URL at it, no certificates needed. cost-xray does exactly that. I grepped the whole repo for TOOL_SEARCH. Nothing.

So under cost-xray, tool search is off, and every MCP schema goes back into the prefix on every turn. The “unused MCP waste” panel then shows real waste, but waste that was mostly deferred before you installed the tool. It is like checking tyre pressure with a gauge that lets a little air out every time.

Two knock-on effects follow. The prefix is bigger, so every warm read and every cold rewrite costs more. And with schemas in the prefix, connecting or removing an MCP server mid-session breaks the cache again, where with deferred tools it would only have appended a line.

This is not a cost-xray quirk. On 16 September an issue on OpenCodex reported the same thing through its proxy: 292 tool schemas loaded up front, 248,600 tokens of MCP definitions, more than a 200k window can hold. With tool search on, the same schemas came to “a few hundred tokens.” On 25 September the distil project merged a fix for its own wrapper. Its note says that across 14,253 anonymised requests, MCP definitions were 7.6% to 24.5% of the bill. cost-xray’s last commit predates both.

The fairest defence: caching still works through cost-xray. mitmproxy forwards the cache markers untouched, so a bloated prefix is mostly billed at $0.20, not $4. On a warm day the extra is cents. The damage shows on cold turns and in window space, and those are exactly what the tool is meant to diagnose.

The fix is one variable. Put it in your shell, or in the env block of ~/.claude/settings.json:

export ENABLE_TOOL_SEARCH=true

One caution. A March issue on the Claude Code tracker found the flag did not stick behind a proxy, because the capability check read an internal Haiku call. OpenCodex reports the manual flag working now. Don’t take either on faith. Start claude --debug, search the output for “Tool search disabled,” and watch the MCP rows in cost-xray shrink to names.

Two more cost-xray gotchas I found in the code

Prices come from LiteLLM’s public price list, fetched from raw.githubusercontent.com once a day. If that fetch fails, say on a locked-down work laptop, cost-xray falls back to a snapshot bundled in the repo. That snapshot has no Opus 5.5 or Sonnet 5, so both get the default: $5 input, $25 output, $0.50 cache read. I forced the fallback and re-ran the same turn. The warm turn jumped from $0.0088 to $0.0206, cache reads overstated 2.5 times. The cold turn went from $0.3066 to $0.3833.

Token totals stay right, because cost-xray scales its estimates to the usage numbers Anthropic returns. Only the dollars drift. Check that ~/.cost-xray/litellm_prices.json exists after your first run. If it does not, allow that one host.

The tokenizer has the same silent fallback. My sandbox blocked the download of OpenAI’s o200k_base vocabulary, and cost-xray quietly dropped to “characters divided by four” for sizing each piece. The split between sources gets rougher. No warning appears.

Last, the window bar. cost-xray’s Claude Code adapter assumes a 200,000-token window unless the request carries a context-1m beta header or a [1m] model tag. Opus 5.5 defaults to a 1M window. If your bar looks five times more crowded than /context says, that is why.

How to set up cost-xray in 10 minutes without skewing the numbers

Use a throwaway repo, not the one you get paged for. cost-xray runs on macOS and Linux. The project is at tigerless-labs/cost-xray.

curl -fsSL https://raw.githubusercontent.com/tigerless-labs/cost-xray/master/install.sh | bash
echo 'export ENABLE_TOOL_SEARCH=true' >> ~/.bashrc   # or ~/.zshrc

Pick Claude Code when the installer asks. Open a new terminal, because the wrapper only applies to shells started after install. Run claude and do one real task, something that reads files and runs a test. Then type cx from any folder.

Drill down from agent to project to session to category to MCP server. With tool search working, idle servers should be nearly invisible. Then open /usage in Claude Code and compare the cache hit rate with what cost-xray shows. They should roughly agree. If cache-write dollars keep climbing on the same source turn after turn, something above it in the prefix keeps changing. That is your leak.

When you are done, cx stop pauses capture and cx uninstall removes the wrapper and services. Your captures stay in ~/.cost-xray/ until you delete that folder.

So where does the Claude Code bill actually go now?

It goes to misses, long contexts and thinking tokens. Idle schemas come a distant fourth, as long as tool search is on.

That changes the advice. Pick your model and effort level at the start of a session and leave them. On Opus 5.5 with an API key or subscription, changing effort keeps the cache, but switching models never does. Don’t connect or disconnect MCP servers mid-task. Use /clear between unrelated jobs, and /rewind rather than /compact when you just want to back out of a wrong turn, since rewinding lands on a prefix that is already cached.

If you are on an API key and you step away a lot, the one-hour cache can pay for itself. The maths: a one-hour write costs $3 more per million new tokens than a five-minute write. One avoided miss on a 150,000-token session saves $0.75. If each turn adds about 3,000 new tokens, the premium is under a cent a turn, so one long break every 80 turns covers it. Set promptCacheTtl to 1h and see.

The bigger lesson is about measuring anything. A meter that changes the circuit is not a meter. cost-xray is the closest thing to an X-ray of a Claude Code request that exists right now. Set one variable before you trust the picture.

Where this fits with the other Claude Code notes

This is the follow-up to the AutoHarness post, which ended on the line “the pain is a rewrite, and not being able to see which block it was.” cost-xray is how you see the block. If you were about to raise Opus 5.5 effort to rescue a sluggish session, read the prompt audit first. More effort means more thinking tokens, and those bill at the output rate.

A bloated CLAUDE.md sits in the prefix of every request too, so the 65-line version is also a cost fix. And if you want to run cost-xray somewhere it cannot touch your real machine, the coop sandbox walkthrough sets up exactly that.

Common questions about Claude Code MCP token cost

Do unused MCP servers still cost tokens in Claude Code?

Very little, on current versions with tool search on. Only tool names and server instructions enter the request until Claude uses a tool. They cost full schemas on every turn when tool search is off, which happens behind a custom ANTHROPIC_BASE_URL, on Amazon Bedrock, Google Cloud or Microsoft Foundry, or when ENABLE_TOOL_SEARCH=false. See also the Claude Code MCP docs.

Why did my Claude Code bill jump after a break?

Your cache expired and the next message rewrote the whole prefix. On an API key the main conversation defaults to a five-minute cache, and on Opus 5.5 a rewrite costs 25 times a cached read. Setting promptCacheTtl to 1h helps if you often step away.

Is cost-xray safe to run on my work machine?

It binds to 127.0.0.1, sends no telemetry, and strips Authorization headers, API keys and cookies before writing to disk. It still stores your prompts and file contents under ~/.cost-xray/. Treat that folder like your shell history, and check your company’s policy first.

What is the difference between cost-xray and ccusage?

ccusage reads Claude Code’s local transcripts and totals cost by session, day or model. cost-xray reads the actual API request, so it can price the system prompt, each tool schema and each MCP server, which never appear in transcripts. Use ccusage for “how much,” cost-xray for “why.”

If you only do one thing today, run /usage and read the Prompt cache line. If the misses surprise you, install cost-xray with tool search on and find the block that keeps changing.

JOIN OUR NEWSLETTER
Be the first to know. Get fresh AI/Tech updates instantly, no spam, unsubscribe anytime

2 comments

Leave a comment