RTK vs Caveman: Why Big Token Savings Shrink in Real Claude Code Work

2026-09-30

Real tests. Plain words. No hype.

TL;DR: RTK shrinks the terminal output Claude reads. Caveman shrinks the words Claude writes back. Both do what they say, but in real Claude Code sessions each cuts only a small slice of the bill, often under 10%, because most of what you pay for is Claude re-reading a conversation it already has.

You run Claude Code for an hour and hit your usage limit. Or the API bill lands and it is bigger than you expected. Then a viral list promises that two free tools, RTK and Caveman, will cut your tokens by 60 to 90%.

This post shows what each tool actually touches, why that headline number shrinks in real coding work, and gives you a short script to check the math with your own numbers.

What is a token, and why does Claude Code burn so many?

Simple mind map of RTK vs Caveman token savings in Claude Code
RTK vs Caveman at a glance: what each tool touches and why savings shrink.

Every turn, Claude re-reads the whole conversation, so tokens pile up fast.

A token (a small chunk of text, roughly three quarters of a word) is the unit you pay for. Claude reads tokens in and writes tokens out, and both are billed.

Claude Code works in a loop. It reads your request, calls a tool (runs a command, opens a file), reads the result, writes a reply, then goes again. Each lap of that loop is a turn.

Here is the part most people miss. The model has no memory between turns. So on every turn, Claude Code sends the entire conversation again: the system prompt (the long set of built-in instructions Claude Code starts with), your messages, and every tool result so far.

Think of it like re-reading a whole notebook from page one every time you add a sentence. By turn 30, you are paying to re-read 29 turns of history.

Claude Code agent loop re-sending the whole conversation every turn
The agent loop re-sends the whole conversation on every turn.

Why re-reading does not cost full price

Prompt caching makes re-reads cheap, and that quietly decides which savings matter.

Prompt caching (the API keeps a copy of text it saw recently and charges much less to read it again) changes the bill a lot. Instead of one price per token, you pay three different prices.

On Claude Opus 5.5 the list prices per million tokens are: $5 when new text first enters the cache (a cache write), $0.20 when old text is re-read from cache (a cache read), and $20 for every token Claude writes back (output). Plain uncached input is $4.

Claude Opus 5.5 prices for cache reads, cache writes, input and output
Cache reads cost a hundredth of output tokens on Claude Opus 5.5.

Two things follow from that. A token that enters the conversation early gets re-read on every later turn, so cutting it early saves a little on every turn after. And a word Claude writes costs 100 times more than a re-read, but Claude writes far fewer words than it reads.

Keep those two facts in mind. They explain everything below.

What does RTK actually do?

RTK filters shell command output before Claude sees it, but only for commands that run through Bash.

RTK is a CLI proxy (a small program that sits in front of another command and trims its output). You run rtk init -g once, and it adds a hook (a script Claude Code runs automatically before a tool call) that quietly rewrites commands like git status into rtk git status.

The filters are smart. Test runs keep only the failures, git diff loses its noisy headers, and grep results get grouped by file. Running rtk --help shows filters for git, test runners, cargo, npm, docker, kubectl, and more, plus rtk gain (a running savings tally) and rtk recall (brings back output a filter hid, so nothing is lost for good).

The catch is reach. Claude Code reads files with its own built-in Read, Grep and Glob tools, and those never go through Bash, so RTK never sees them. RTK’s own docs say this. On top of that, many shell commands agents run, like python3 scripts, have no RTK filter.

In a paired test on 86 real coding tasks, the commands RTK could touch carried just under 20% of all the tool output in the session. RTK can squeeze that slice hard. It just cannot touch the other 80%.

RTK hook reaches about 20 percent of Claude Code tool output
RTK only reaches about a fifth of tool output in a Claude Code session.

What does Caveman actually do?

Caveman makes Claude’s prose short and leaves code, commands and error text alone.

Caveman is a skill (a markdown instruction file Claude Code loads to change how it behaves). It tells Claude to drop articles, filler words and pleasantries, and to keep every technical word.

It has sensible guardrails. It never drops words like “not”, “never” or “only”, because losing them flips the meaning. It switches back to full sentences for security warnings and for confirming anything you cannot undo. It does not touch code comments, commit messages or docs.

The famous 65% comes from chat-style answers, where almost everything Claude writes is prose. In a coding session, most of Claude’s output is code, diffs, tool calls and exact error strings, and Caveman correctly leaves all of that as is.

So the real number is smaller. Across 82 paired coding tasks, Caveman cut output tokens by 8.5%, with no measurable drop in quality.

Caveman shortens prose but leaves code and tool calls unchanged
Caveman shortens prose but leaves code and tool calls untouched.

RTK vs Caveman: which one saves more tokens?

On paper, RTK wins in coding sessions and Caveman wins in chat, and neither gets near the headline number.

To see why, I wrote a small model of one session and ran it. It assumes 50 turns, a 20,000-token starting prompt, about 2,000 tokens of tool output and 400 tokens of reply per turn, with a quarter of each reply being prose. Prices are Claude Opus 5.5 list prices.

SetupSession costCheaper by
No tools$1.890%
RTK (reaches 20% of tool output, cuts 80% of it)$1.738.4%
Caveman (cuts 65% of the prose)$1.795.1%
Both together$1.6313.5%
RTK if it could reach all tool output$1.1041.9%
Caveman in a chat-only sessionn/a57.3%

Three lessons fall out of this. RTK’s limit is reach, not compression: give it all the tool output and the saving jumps to about 42%. Caveman really does shine, but in chat, not in coding. And these numbers are a best case, as the next section shows.

Modeled bill savings for RTK, Caveman and both together
Modeled savings for RTK, Caveman and both together.

Want to check with your own numbers? Save this as token_bill.py, change the values in session(), and run it with python3 token_bill.py.

PRICE = {"cache_read": 0.20, "cache_write": 5.00, "output": 20.00}  # Opus 5.5, $ per million

def session(turns=50, system=20_000, tool=2_000, out=400, prose=0.25,
            tool_cut=0.0, prose_cut=0.0):
    ctx, bill = system, {"cache_read": 0, "cache_write": system, "output": 0}
    for _ in range(turns):
        o = out * (1 - prose * prose_cut)   # Caveman trims only the prose part
        t = tool * (1 - tool_cut)           # RTK trims only the output it can reach
        bill["cache_read"] += ctx           # whole history re-read from cache
        bill["output"] += o
        bill["cache_write"] += o + t        # reply + next tool result join the cache
        ctx += o + t
    return sum(v * PRICE[k] / 1e6 for k, v in bill.items())

base = session()
for name, kw in [("RTK", dict(tool_cut=0.16)), ("Caveman", dict(prose_cut=0.65)),
                 ("Both", dict(tool_cut=0.16, prose_cut=0.65))]:
    cost = session(**kw)
    print(f"{name:8} ${cost:.2f}  {100*(base-cost)/base:.1f}% cheaper")

It is a simple model, not a bill. It ignores cache expiry and uses round numbers. But it shows which lever moves which part of the cost.

Why real sessions save even less than the math

Shorter output can change what Claude does next, and one extra turn wipes out the savings.

When Claude sees less, it sometimes goes back and looks again. Every extra look is a full turn, and a full turn means re-reading the whole conversation.

That is exactly what showed up in the paired RTK test. At low reasoning effort, sessions with RTK took about 13.8% more turns and cost a median 7.6% more per task. At high effort the difference vanished. Task quality was the same either way.

There is a second trap. rtk gain compares filtered output with raw output, and it reported 99.8% saved in that same test while the real bill went up. Claude Code already cuts huge tool output short on its own, so the scary “10,000-line log” was never billed in full. The counter measures text that was never going to cost you.

Should you install RTK or Caveman?

Install them for cleaner, longer sessions, not to halve your bill.

  • Lots of test, build and log output through Bash: try RTK. Less noise in context means sessions last longer before Claude has to compact (summarize old history to free up space).
  • Mostly chat, planning and questions: try Caveman. Prose is exactly what it cuts.
  • You want a lower bill: measure first. Run the same kind of task with and without the tool, and compare your billed cost, not the tool’s own counter.
Decision guide for choosing RTK or Caveman
Which one to install, based on how you use Claude Code.

To try them, install RTK with brew install rtk and run rtk init -g. Add Caveman with npx skills add JuliusBrussee/caveman -g, and turn it off any time by typing “stop caveman” or “normal mode”.

If your real problem is how much Claude Code sends before you even type, look at your setup before adding more tools. The CLAUDE.md file that fixes AI coding shows how to keep project instructions lean, and running Claude Code and Codex together covers splitting work across agents. For running agents safely in a box, see the coop Claude Code sandbox.

Links: RTK on GitHub · Caveman on GitHub · Claude pricing docs · Full paired benchmarks: RTK, Caveman

Common questions about RTK vs Caveman

Does Caveman make Claude worse at coding?

In a paired test on 82 real coding tasks, quality was statistically the same with and without it: 8 tasks improved, 10 got worse, 64 tied.

Does RTK work with Claude Code’s Read and Grep tools?

No. RTK hooks into Bash commands only, and Claude Code’s built-in Read, Grep and Glob tools skip that hook.

Can I use RTK and Caveman together?

Yes. They touch different sides of the bill: RTK trims command output going in, Caveman trims Claude’s prose coming out. On paper, the two together cut about 13.5% in a coding session.

Why does rtk gain show huge savings when my bill barely moved?

It compares filtered output with raw output, including huge output Claude Code would have cut short anyway, and it cannot see extra turns. Compare your billed cost instead.

Real tests. Plain words. No hype.

JOIN OUR NEWSLETTER
Be the first to know. Get fresh AI/Tech updates instantly, no spam, unsubscribe anytime

3 comments

Leave a comment