SoL-Pi’s 49% token cut is the setting that drops the score

2026-09-26

The number going around today is 49 percent. Nvidia’s SoL-Pi, people say, cut a coding agent’s tokens almost in half and the model was not the expensive part. A few posts round that to 50 percent versus Codex, or to $13.50 an hour.

I read arXiv 2609.20519v1 (posted 17 September 2026), the README and docs/configuration.md on NVlabs/SoL-Pi, and the TypeScript those docs point at. Then I cloned main and ran the test suite. The 49 percent is in the paper. It is the four-switch stack, on GPT-5.6 Sol, against Pi, on EdgeBench. That same row scores 42.003. Pi, with no SoL-Pi switches on, scores 44.833.

If you install the repo and do nothing else, you get none of this. Every mechanism ships off. I did not re-run EdgeBench. The efficiency row in the paper cost the authors $894 in API calls at 17 August 2026 prices. I am not going to invent a transcript for that.

The 49% is one cell, and it is not the cell versus Codex

Mind map of SoL-Pi token cut: four switches, ObservationPack vs all-four scores, install defaults off, EdgeBench tradeoffs

EdgeBench here is 51 public tasks. The search that invented the mechanisms was not allowed to see them. Eleven tasks were a one-way check after a candidate was frozen. Forty were the final look. Held-out results did not go back into the search. That part of the setup is better than the usual “we tuned on the test.”

Recorded token traffic in their tables is input plus cache read plus cache write plus output. Not “output tokens.” Not “the expensive ones.”

HarnessModelTokens (B)API costAvg. score$ per score point
CodexGPT-5.6 Sol3.0537$1,78734.7381.0086
PiGPT-5.6 Sol2.1538$1,33944.8330.5855
SoL-Pi, all fourGPT-5.6 Sol1.0990$89442.0030.4174
SoL-Pi, ObservationPack onlyGPT-5.6 Sol2.0224$1,27147.2080.5280
Claude CodeOpus 52.0045$2,53543.6891.1377
PiOpus 52.3697$1,74144.7560.7625
SoL-Pi, all fourOpus 51.3101$1,15842.2240.5376
SoL-Pi, Action Fusion onlyOpus 52.1016$1,60550.4820.6235

Those are Table 2’s point estimates. The paper says API prices are the ones from 17 August 2026, and that prices move. I am not re-pricing this on a September card.

On Sol, all four mechanisms versus Pi: tokens 2.1538 to 1.0990 is a 49.0 percent cut. Score 44.833 to 42.003 keeps 93.7 percent. Dollars $1,339 to $894 is 33.2 percent off. The abstract calls that “comparable.” You can call a 2.8 point dip comparable. You should not call it the same score.

Same stack, moved to Opus 5 with no extra search: tokens down 44.7 percent (2.3697 to 1.3101), score 42.224 against Pi’s 44.756 (94.3 percent), cost down 33.5 percent ($1,741 to $1,158). The authors say the switches fired less often on Opus, because the search only ever watched Sol trajectories.

“50% versus Codex” is a dollar, and Codex was already behind

Figure 1 in the paper says the full stack cuts API cost 50.0 percent versus Codex on Sol, and 54.3 percent versus Claude Code on Opus. Check the subtraction. ($1,787 − $894) / $1,787 is 50.0 percent. ($2,535 − $1,158) / $2,535 is 54.3 percent. Cost. Not tokens.

Tokens versus Codex fell further than the dollars: 3.0537 to 1.0990 is about 64 percent, because Codex’s bulk is cheap cache reads (3.0287 of that 3.0537). Tokens versus Claude Code fell less than the dollars: 2.0045 to 1.3101 is about 35 percent. Claude Code’s cache-write column on that Opus run is 0.1701 billion tokens. Pi’s is 0.0348. The August price card punished the rewrite, so the bill looks worse than the token total.

Codex on this bench scored 34.738. The four-switch stack scored 42.003 and cost half as much. That is a real win against Codex. It is a different sentence from “half the tokens, same work as the harness that already led the table.” Pi was the score leader before SoL-Pi existed, at $1,339 and 44.833.

On the same Sol block, OpenCode is $3,422 for a 29.552, with cache writes of 0.2865 billion tokens. Oh-My-OpenCode scores 38.523 at $2,678. A harness can spend more than the model choice. That is the useful half of the pitch. It does not require the 49 percent.

The abstract also prints hourly savings of $8.75 to $13.50 versus Codex and Claude Code, and $4.36 to $5.71 versus Pi. I did not find the clock next to Table 2. Dividing the Codex gap by the “official @2h” label is about $446 an hour, which is not $13.50. So that range is some other burn rate. Quote it as their estimate.

The row almost nobody pasted beats Pi

SoL-Pi [Performance] in Table 2 is not a mystery blend. The caption says it is the single mechanism with the best score for that model. ObservationPack on Sol. Action Fusion on Opus.

On Sol, ObservationPack alone scores 47.208 against Pi’s 44.833. That is up 5.3 percent, on 2.0224 billion tokens instead of 2.1538 (about 6 percent fewer), at $1,271 instead of $1,339. On Opus, Action Fusion alone scores 50.482 against 44.756. Up 12.8 percent, tokens 2.3697 to 2.1016, cost $1,741 to $1,605.

If you turned all four on because a screenshot said 49 percent, you bought the row that gives points back. If you wanted the higher score, you wanted one switch, and which switch depends on the model. I did not re-benchmark that. I am reading their caption.

Cache writes on the four-switch Sol run go from 0.0141 billion to 0.0316 billion. Cache reads fall from 2.1326 to 1.0605. Shorter context means the prefix gets rewritten. The bill still drops, at August prices, because the reads were the mountain. A price card that charges writes harder than that August card can shrink the dollar win without touching the token win.

Installing it changes nothing until sol-pi.json says so

The package on main is sol-pi 0.1.0, MIT, a Pi extension. It does not patch Pi and it does not vendor it. Requirements in the README: Node.js 22.19 or newer, and @earendil-works/pi-coding-agent 0.85.1.

npm install --global @earendil-works/pi-coding-agent@0.85.1
pi install git:github.com/NVlabs/SoL-Pi

Project-local, which is what I would use:

pi install git:github.com/NVlabs/SoL-Pi --local --approve

Config is one file. Search order: .pi/sol-pi.json in the project, only if Pi has marked the project trusted, otherwise ~/.pi/agent/sol-pi.json. If neither exists, built-in defaults. The project file replaces the global file. They do not merge. A global observationPack: true disappears the moment a project file exists and only sets actionFusion.

DEFAULT_CONFIG in src/sol-pi/config.ts has all four booleans false. cacheWriteReadRatio defaults to 12.5. That number is not your invoice. Online Context Compact uses it to decide whether a compaction’s cache rewrite is worth it. Zero means “treat a cache write as free,” which makes compaction more willing, not cheaper by magic.

Unknown keys abort the load. The test even checks the typo: actionFussion throws Unknown SoL-Pi config key: actionFussion. "yes" instead of true throws. Version must be 1.

I cloned main at 1559b5c (22 September 2026, “Merge pull request #35” for Windows shell paths in Action Fusion). GitHub’s API said 3,094 stars when I asked, on 26 September. The repo pushed_at that afternoon was not this commit. Main’s tip was still the 22nd. On Node v22.23.3, npm ci --ignore-scripts and npm test (vitest 4.1.9): 19 files, 158 tests, all passed, about 13 seconds. That includes tests/config.test.ts, which checks the disabled default, the no-merge rule, and the untrusted-project rule. One integration test loads the package and runs fused tools inside Pi. I did not open an interactive Pi session, and I did not point it at a paid model.

The Decoder’s writeup is dated 26 September, which is why the percent is moving now. The preprint is nine days older. This is not a repo that appeared this morning.

Four switches, and only one of them calls another model

Action Fusion replaces Pi’s edit and write. An edit and the validation command that was going to follow it can share one tool call, so you skip a model round trip. The paper is explicit that a command which needs to look at the edit first stays separate. Shell behavior stays Pi’s.

ObservationPack is local. In observation.ts, THRESHOLD_BYTES is 10 * 1024 and FULL_SENDS is 2. A tool result over 10 KiB is archived under the session directory and sent whole for the first two provider requests. After that the model sees a stable obs_ handle, the original size, and a 1,024-byte excerpt split into head and tail. Exact pages come back through obs_recall. Smaller results are left alone. The archive is not deleted when the session ends.

Online Context Compact hooks update_plan. When a plan step finishes, it estimates how many requests are left and compares the input you would skip with the cost of rewriting the cache, using that 12.5 ratio, or it compacts when the window is simply full. Compaction is Pi’s own. After it works, the extension sends one hidden generic message with triggerTurn so the same Pi process keeps going. Kill Pi, or cancel, and nothing resumes you. The plan text lands in Pi’s session log. SECURITY.md says to treat that like the rest of the chat. There is no separate sidecar.

Evidence-Preserving Reducer is the one that leaves the machine. Off unless the flag is true. Default route in source, built by joining the pieces so a grep for the product name misses it: provider openai-codex, model gpt-5.6-luna. Same model the paper used, at high in the experiment. It only looks at build and test logs. File reads and search results skip it. The command has to match a fixed list: pytest, npm/pnpm/yarn test, cargo, go test, make, ninja, cmake, lean, lake, coq, zig, ctest, bazel, and a few cousins. Under 4,096 bytes, it does not bother (DEFAULT_MIN_BYTES). A cheaper model writes a receipt. A checker then demands the schema, the source hash, the exit status, exact quotes, and a smaller size. Fail, or a suspected credential, or a receipt that is not smaller, and the original log stays. The secret check is one regex (api_key, bearer, access_token, secret, then an = or :). SECURITY.md says that detector is a precaution, not a scanner. Do not enable this for logs that have to stay local. Timeout in source is 90 seconds. The receipt caps at 12 evidence items and 600 characters a quote. If Luna is not in Pi’s model registry, the original tool result is what the agent sees. Credentials do not go in sol-pi.json. Auth is Pi’s.

The reducer runs before ObservationPack. A receipt is marked so ObservationPack will not pack it again.

It also drops three solved tasks

Terminal-Bench 4, 63 CPU-only tasks, GPU tasks left out: Codex and Pi each solve 18. SoL-Pi solves 15. Total model cost $211.12 versus Pi’s $286.45 (down 26.3 percent) and versus Codex’s $272.35. Cost per solved task: $14.07 versus Pi’s $15.91 and Codex’s $15.13. Cheaper, and three fewer solves. “Generalizes” in the paper means the bill moved. It does not mean the solve count held.

IMO 2026, six problems, GPT-5.6 Sol at xhigh, answers checked in Lean 4: SoL-Pi passes 3, same as Pi, at $62.69 total and $20.90 per pass. Pi was $75.95 and $25.32 per pass. Codex passes 5 at $114.47, $22.89 each. Lowest dollars per pass is not the same as most problems solved. If the job is “get the proof,” Codex won that table. If the job is “stop paying Pi’s overhead on the three you can already get,” SoL-Pi did that.

In the kernel swarm (a coordinator plus workers, two-hour chart): SoL-Pi workers, 1,127 cycles at $60.11. Pi workers, 1,366 cycles at $82.12. A single Codex agent, 1,333 cycles at $39.20. The SoL-Pi swarm and the single agent cleared every speed threshold. The Pi swarm missed the last one, under 1,363 cycles. The cheapest line is still one agent. The paper does not claim a swarm is how you save money. It claims a cheaper worker harness spends a fixed swarm budget better than Pi’s workers. A Codex developer, Eric Provencher, has been saying more than two sub-agents usually burn the budget checking each other. I would not turn this paper into a reason to spawn twenty of them.

What I would actually put in the file

Not on the repo you get paged for. Pi 0.85.1, Node 22.19 or newer, project trusted, or the .pi/sol-pi.json is ignored and you will think the extension is broken.

The README’s conservative config is the two local switches. No second model call, and compaction stays off so a run is not cut and resumed:

{
  "version": 1,
  "actionFusion": true,
  "observationPack": true,
  "evidencePreservingReducer": false,
  "onlineContextCompact": false,
  "cacheWriteReadRatio": 12.5
}

That is not the 49 percent row, and it is not the Table 2 “performance” row either. It is the “don’t phone home, don’t compact out from under yourself” row. I have not measured its score.

If you are copying the paper’s best single switch and you are on GPT-5.6 Sol, the caption says ObservationPack only. On Opus 5, Action Fusion only. Leave the other booleans off rather than “true, why not.” All four is how you get 42.0 instead of 47.2 on Sol.

Turn the reducer on only after you have read SECURITY.md and you are fine with pytest and npm test output going to openai-codex / gpt-5.6-luna through Pi’s auth. Archives land in <session-directory>/sol-pi/<session-id>/, and they stay after you quit.

Skip it entirely if you wanted a Claude Code plugin. The Claude Code column is a comparison harness in the paper, not an install target. Skip it if you cannot pin Pi to 0.85.1. Skip the hourly dollar quote until you have an hour of your own traces. And if a post says the model was not the bottleneck, ask which row. On this bench the model was held still, and the harness still moved the score by more than the token cut moved the brag.

Where this sits next to the other harness notes

SoL-Pi is a bill for how much text gets replayed, not a new model. The same split is why Claude Code’s startup text and MiniMax Code’s are different sizes before you type. If a session feels dumb and the instinct is to raise Opus 5.5 effort, the 23 September prompt audit still comes first. Effort multiplies work. It does not stop a 10 KiB log from being pasted back on every turn.

A skill file is a different layer. AutoHarness only archives skills it wrote. SoL-Pi will not prune those, and AutoHarness will not fuse your edit with the test command. The map of how Claude Code is wired is the right background if “harness” is still a vague word. The short CLAUDE.md a lot of people pinned changes what the model is allowed to do. SoL-Pi changes how often Pi pays to look at what already happened.

After this ships, the link belongs under that MiniMax comparison, on the sentence about harness overhead. People land there wanting a smaller startup prompt. This is the follow-up for the long run, where the log is the bill.

Common questions about SoL-Pi token cut

Does SoL-Pi cut my Claude Code bill?

No. It is an extension for Pi 0.85.1. Claude Code in the paper is the baseline they measured, not a plugin you enable. On Opus 5 that baseline cost $2,535 for a 43.689. Pi plus all four SoL-Pi switches cost $1,158 for a 42.224. Different product.

Which switch is the 49 percent?

All four, on GPT-5.6 Sol, versus Pi, on EdgeBench. Score goes from 44.833 to 42.003. One switch, ObservationPack, is the Sol row that scores 47.208. On Opus 5 the single switch that scores 50.482 is Action Fusion, not the four-stack.

Will it send my test logs to another model?

Only if evidencePreservingReducer is true. The default route is openai-codex / gpt-5.6-luna. It only fires for a fixed list of build and test commands, and only at 4 KiB or more. File reads are skipped. A short regex drops lines that look like tokens. That is not a guarantee. Leave the flag false if the log has to stay on the machine.

Why did install do nothing?

Defaults are off. No sol-pi.json, no behavior. A project file is ignored until Pi trusts the project. A typo in a key, or "yes" instead of true, stops the extension from loading. Project config does not merge with the global file.

If you try one thing, pin Pi to 0.85.1, trust a scratch repo, and paste the two-switch file from the README. Then read one archived tool result under sol-pi/ before you decide the handle is enough. A 49 percent quote that does not name the score next to it is a different product from the one in Table 2.

JOIN OUR NEWSLETTER
Be the first to know. Get fresh AI/Tech updates instantly, no spam, unsubscribe anytime

Leave a comment