AutoHarness will retire skills. Just not the ones you wrote.

2026-09-26

4,124 people have starred a Claude Code plugin whose pitch is that it cleans up after itself. Trendshift had it at number one this morning, momentum 139, ahead of Paperclip and Orca. Four verbs: learn, merge, update, prune.

Here is the part the one-liner skips. AutoHarness archives only the Claude Code skills it created itself, and only after they go unused. Skills you wrote, and skills you installed from a marketplace, stay on disk. I checked that guard in v0.5.3, then ran the lifecycle tests that pick who gets archived.

If you wanted a janitor for the 2,000 skills you installed and forgot, this is the wrong tool. If you are tired of writing the same lesson down by hand and watching near-copies pile up, that is the job it actually does.

Why unused Claude Code skills still sit on the bill

Mind map of how AutoHarness prunes unused Claude Code skills it created, from reflect to archive

Claude Code does not send “your repo” as a feeling. It sends text, and you pay by how much text goes in and how much comes back. That unit is a token (a chunk of a word the bill counts as one). The amount of text the model can hold at once is the context window (the desk it is working on, and it fills up).

A skill is not a plugin and not a prompt you paste each morning. It is a folder with a SKILL.md file (a short instruction sheet the agent can pull when a task matches the description at the top).

Current Claude Code notes, from the 2.1.280 era, say those descriptions load before you type. The bodies load when a skill is invoked. After a compaction, invoked bodies can come back, capped at 5,000 tokens per skill and 25,000 total. “Every skill loads” is true for the description line. It is not true for the whole manual.

An older measurement is easy to misread and still useful. On 12 May 2026 a post in r/ClaudeAI reported 2,596 installed skills, 40 ever invoked, and 102,651 tokens in every session. That is about 40 tokens a skill: a description, not a textbook. The author priced it near $91 a month on the rates they were paying then. I am not re-billing that on today’s card. Opus 5.5 cache reads are $0.20 per million tokens, so a cached 100,000-token block is about two cents a turn. The pain is a rewrite, and not being able to see which block it was.

On 29 August 2026, XDA opened a fresh session and found 51,400 tokens already taken. System prompt about 10,700. Built-in tools about 28,500. The rest was stuff they had installed.

None of that pile is what AutoHarness deletes. The Trendshift blurb invites exactly that mistake.

Four GitHub repos share the name

Search autoharness and you get a pile, not a product.

bes-dev/autoharness builds a project harness from a task description, then patches it after the session. shintaro-amaike/autoharness-skill turns a DeepMind paper into lint rules and a check script. redker56/auto-harness, with the hyphen, runs planning and QA files under .harness/. Tigerless Labs is the one Trendshift ranked first. MIT, created 9 June 2026, last push 25 September 2026, 4,124 stars and 304 forks on the API I pulled. It does not do those other jobs.

The README opens with a number that is not this repo’s score. They link a paper, HAL, where the same model moves from 42% to 78% on CORE-Bench when the surrounding setup changes. The borrowed point is that the harness (the loop, the tools, and the files around the model, separate from the model itself) does a lot of the work. AutoHarness bets one slice of that harness, the skill folder, can maintain itself. They also say they do not spend tokens on a held-out benchmark. If someone quotes 78% as their result, they did not read past the first paragraph.

Rising Repo’s 25 September table, in the slice I went through, did not show this repo among the sharpest star-gain rows. Mentions are hot. The star chart is not exploding this week. Early for a careful explainer. Late for “nobody has heard of it.”

How AutoHarness prunes unused Claude Code skills

A hook counts tool calls. It does not grade the work. The default is 50. The turn that crosses 50 ends with a background reflection, and your session is not blocked. A chat where you only talk never trips it. Talk is not a tool call.

That reflection is a reflector (a child session that proposes add, merge, patch, or drop, and is not given write tools). The promoter (the only part allowed to save a skill) checks the proposal in memory: safety, shape, a reason, a slice of evidence. Fail that and nothing is written. The next session starts with one line on what landed and what was rejected, then that note file is deleted so it does not nag twice. The durable copy stays in the state folder.

Project skills land in .claude/skills/. Habits that are not about one repo land in ~/.claude/skills/. Beside SKILL.md you get a ledger (an append-only .ledger.jsonl that records why the skill was born or changed, pointing at a redacted slice of the session). You also get .sidecar.json. That file is the permission system.

It does not always add. Same scenario, it is supposed to patch the old skill instead of stacking a near-copy. A new lesson that contradicts an old one has to rewrite the stale skill in the same run. A merge must name the skill that absorbed the one that left. An invented umbrella name fails the whole proposal.

Retirement is not a delete. The folder moves to .claude/skills/.archive/<name>/, ledger included. Move it back and it is alive again.

Who moves is evaluate in lifecycle.py. A new skill is on probation (a waiting period counted in requests, not days, so a closed laptop does not age it out). The project layer waits 100 requests. The global layer waits 300, because a global skill shows up in every project. During probation it can still be recalled. It cannot be archived, and it does not count against the shelf cap.

After that, two deaths, and people mix them up.

If it was never loaded and never even looked at, graduation review archives it even when the shelf has room. A look is a view (a session that read into the skill’s folder). A view is not a use (the model actually invoking the skill). A view spares you at graduation. It does not help you later.

If it has been used at least once, low usage will not kill it. It dies only when the mature pool is over the cap, and the lowest rate goes first. Project cap 50. Global cap 20. The rate is uses divided by requests since the skill was created.

The session-start index is bounded by those caps. One line per live skill, description cut at 60 characters. Fill both caps and you get 70 short lines, not 2,596 marketplace blurbs. AUTOHARNESS_INDEX_SUSPENDED=1 turns the index off and leaves the rest running.

A slower curator pass, default every 250 tool calls, folds near-duplicates across the whole library. It keeps 5 snapshots first. A merge is the one edit a single rename cannot undo.

Can AutoHarness delete skills I didn’t write?

No. Not in v0.5.3, which is what .claude-plugin/plugin.json says.

Session start builds its list in _members. For every SKILL.md, it asks sidecar.is_agent_created. That reads .sidecar.json and returns true only when created_by is agent. I called it against a directory that does not exist. It returned false. A skill you wrote by hand, with no sidecar, never enters the archive list and never enters the injected index.

The promoter has a second lock. In validate.py, a modify is rejected unless the target is already agent-created. The finding is self_produced: “target live skill not created_by:agent”. A reject means zero writes.

So “prunes the ones that stop getting used” is incomplete in the way that matters. It prunes the ones it wrote. Your hand-written skills, your marketplace installs: invisible to this pass.

One edge. The marker is created_by, not “a human edited this later.” If AutoHarness creates a skill and you rewrite the SKILL.md by hand, the sidecar can still say agent, and the promoter may patch that folder again. If you want a skill to be yours alone, write it in a folder that never got that sidecar.

Uninstall stops the hooks. It does not delete the files. State lives in ~/.claude/autoharness/ or <repo>/.claude/autoharness/, plus the skill folders. You delete those yourself. The ledger is how you tell its skills from yours.

I ran the lifecycle tests. Rare is not unused.

I did not point a live Claude Code session at a customer repo. This environment does not have your login, and I am not going to invent a transcript. I ran the decision function the session-start hook calls, using the tests in the repo.

Twelve tests in tests/test_lifecycle.py passed when I imported them and called them. The thirteenth needs a pytest fixture, so I called evaluate myself: review suspended, zero uses, zero views, 20 requests, maturity 10. The archive list came back empty. That is AUTOHARNESS_GRADUATION_SUSPENDED, so you do not punish skills for a broken recall surface.

The cases that define the word prune:

A skill named new, zero calls, 5 requests, maturity 10. Not archived. Still on probation.

A skill named idle, zero calls, 20 requests, maturity 10, capacity 5. Archived, even though the shelf was not full.

A skill named rare, one call, 1,000 requests, capacity 5. It stayed. Low rate is not death until the pool is over the cap.

Viewed 3 times, never invoked: stayed. Never viewed and never used: archived.

Then a capacity fight. viewy has 1 use and 99 views. usey has 50 uses and no views. Capacity 1. viewy is archived. The test is test_view_never_feeds_the_rate. Looking is not using.

Unused gets filed even when there is room. Rarely used stays until the shelf is crowded. That is idle versus rare, and it is the only summary I would bet on.

AutoHarness versus writing the skill yourself

You write itHermes-style timerAutoHarness 0.5.3
What starts learningYou, sitting downIdle time, plus a daemon50 tool calls, or /learn
What it may retireWhatever you deleteIts own library, on a clockOnly created_by: agent, archived
Needs a benchmarkNoNoNo
Needs a resident processNoYesNo
Shows its own list to the modelOnly if the host doesYesYes, capped

Hermes, from Nous Research, is the comparison the README picks. AutoHarness studied that design, then refused the daemon and the wall-clock aging. A skill you have not needed this week is not dead. A skill that had a fair number of requests and was never opened might be.

The cost to watch is the reflection, not the index. Every 50 tool calls it spawns a child session. Those tokens will not show up as a line in your main chat. The README’s fast config sets that number to 3 so you can watch a folder appear. That is a bad daily setting. The carrier is still bundle: a redacted window, not a fork of your warm cache. They leave the fork off until they have measured the cache hit. Reflection today is extra work, not a free reread.

There is a body cap. AUTOHARNESS_SKILL_BODY_MAX_LINES defaults to 25 non-blank lines. Longer than that, the promoter treats it as a transcript and rejects it. Detail belongs under references/. If your skills are essays, this tool will refuse them.

Try this on a throwaway repo tonight

Not on the monorepo you get paged for. You need Python 3.11 or newer on the PATH the Claude Code process can see. Without python3, the hooks never fire. The badge says Linux and macOS. I would not promise Windows.

In the Claude Code box, run /plugin marketplace add tigerless-labs/autoharness, then /plugin install autoharness@autoharness, then /reload-plugins or restart.

For this experiment only, put this in .claude/settings.json:

{ "env": { "AUTOHARNESS_REFLECT_EVERY_N": "3", "AUTOHARNESS_MATURITY_PROJECT": "5", "AUTOHARNESS_CAPACITY_PROJECT": "2" } }

Do a few real tool-using turns, or type /learn the moment you figure a step out. That command lives in skills/learn/SKILL.md. It does not write the file itself. The promoter still has to accept the proposal.

Then ls -la the new folder, because the interesting files are hidden: SKILL.md, .ledger.jsonl, .sidecar.json. You want created_by to say agent. You want a reason and an evidence pointer in the ledger. And you want a skill you wrote yourself, no sidecar, still sitting there after a second session.

Put the knobs back when you are done. Third-party marketplaces do not auto-update unless you enable it. Update the marketplace before claude plugin update autoharness@autoharness, or it may say you are current when you are not. Restart after a version bump. It is a new cached copy, not a hot reload.

Skip it if you wanted a report of marketplace token tax. Those audit tools show the data and delete nothing, which is the right shape for that job. Skip it if you cannot tolerate a background child every 50 tool calls. And read one evidence file before you decide you are fine with redacted session slices living next to the skill.

Where this sits next to the other harness notes

The reason a skill body should stay out of the prompt until the task matches is the same idea as how Claude Code is actually built. If the failure you hate is the agent rewriting half the repo, the 65-line CLAUDE.md a lot of people pinned still does more per minute than a self-writing skill layer. AutoHarness will not install those rules for you.

If you were about to crank Opus 5.5 effort because a session felt dumb, read the prompt audit from 23 September first. Effort multiplies the work. It does not clean the desk. For a harness that is not Claude’s, the 8-minute MiniMax Code CLI test is the adjacent post. This one is a plugin inside Claude Code, not a second CLI.

After this ships, the link belongs on that Claude Code walkthrough, in the skills session. People still land there thinking a skill is a prompt. This is the follow-up that says who is allowed to delete one.

Primary source: tigerless-labs/autoharness (and the HAL paper they cite, arXiv:2510.11977, is not this repo’s score).

Common questions about pruning Claude Code skills

Why does every Claude Code skill load into every session?

The description does. The body does not, on current behavior: bodies load when the skill is invoked, and compaction brings back a capped amount. The May figure of about 40 tokens per installed skill fits a description line. That overhead is real. AutoHarness will not remove it for skills it did not create.

Can I stop Claude Code from keeping skills I never use?

For skills AutoHarness wrote, leave the defaults alone. An unused, unviewed skill is archived after probation, even under the cap. For skills you installed, no. This plugin skips them. You uninstall those, or you stop loading that marketplace.

How do I stop AutoHarness from archiving a skill I might want later?

Archives are folders. Move the directory out of .claude/skills/.archive/. To pause the review while you check whether recall is even showing the skills, set AUTOHARNESS_GRADUATION_SUSPENDED=1. A crowded mature pool can still lose its lowest rate. Anything younger than the request threshold stays on probation.

Is AutoHarness worth it if I already write my own skills?

Only if you want a second pile that maintains itself beside yours, and you will pay a child session every 50 tool calls. If you already keep a short set and delete what you do not use, you do not need it. The person this is for learns a fix in a session and never writes it down. /learn is that moment. The promoter can still reject it. Read the ledger before you trust the file.

If you try one thing, use the scratch repo and the fast knobs, then put the knobs back. Open .ledger.jsonl before you open a pull request. A skill that cannot show you the session slice it came from is a rumor with a filename.

JOIN OUR NEWSLETTER
Be the first to know. Get fresh AI/Tech updates instantly, no spam, unsubscribe anytime

1 comment

Leave a comment