Skip to content

I tested NVIDIA's SoL-Pi pruning on real bugs. It cost 30% more.

On real Hono bugs, SoL-Pi's ObservationPack cost 30% more. No pruning rule I tried beat one habit: never change what the model has already seen.

I tested NVIDIA's SoL-Pi pruning on real bugs. It cost 30% more.

A coding agent sends the whole conversation to the model on every request. The provider caches the start of it, so the part that hasn't changed costs a tenth of the price. That makes one habit worth more than any clever trimming: keep the start of the conversation exactly the same. Engineers call this a stable prefix.

NVIDIA's SoL-Pi is a set of four cost-saving add-ons for the Pi coding agent. One of them, ObservationPack, does the opposite: it rewrites the middle of the conversation to make it smaller. I ported it into Empryo, line for line, and fixed three real bugs 36 times.

How the cache bills youLink to this section

Anthropic bills a cached token at a tenth of the input price. A token written to the one-hour cache costs twenty times that. Anything after the first changed block counts as new and is written again.

How far back the removed result sits decides the size of that bill:

A removal only pays for itself after enough requests: tokens billed again × 19 ÷ tokens removed.

One extra request to fetch something back also wipes out about 43 requests' worth of savings.

What ObservationPack doesLink to this section

Any tool result over 10 KB, like a file the agent read or a test log, rides along in full for two requests. Then it shrinks to a stub.

The testLink to this section

35 of 36 runs fixed their bug. The one miss was an ObservationPack run on the router bug that failed 1 of 644 tests.

The result: two losses and a tieLink to this section

Why it lost: the agent went back for what was removedLink to this section

On the form data bug, ObservationPack stubbed out src/request.ts, a 16.8 KB file, two requests after the agent read it. The agent was still editing that file, so it read it again, in small pieces, over and over. The recall tool doesn't help: it returns the copy from before the agent's own edits.

Each request cost the same on both sides, about 1 cent. There were just more of them, and each prune rewrote the cache.

Every way to prune, comparedLink to this section

Before the Hono runs I tried gentler rules on a synthetic project with 19 planted bugs. None beat the stable prefix.

The only rule that didn't lose removed results so new that the cache hadn't moved past them yet. Each request got 16% cheaper, the agent made 17% more of them, and it came out even. Everything that reached further back cost more.

Their numbers and mineLink to this section

NVIDIA tested ObservationPack on its own before shipping it (write-up, paper). Same idea, opposite result:

NVIDIA's testMy test
AgentPiEmpryo
Tasks11 long EdgeBench tasks3 Hono bug fixes
Big resultsKept in fullTrimmed on arrival
Stub2 KB + 1.5 KB512 B + 512 B
Bill23.6% lower30% higher

Across all four add-ons together, they report 45 to 49% fewer tokens than plain Pi, about a third lower cost, and about 94% of Pi's score. This isn't a claim that SoL-Pi is wrong about Pi:

  • Pi keeps large results in full. Their write-up says a large file "reappeared in every later request" in plain Pi. Empryo trims it the first time, so there's less left to save.
  • Their tasks are long. They note that on shorter tasks, savings "have little time to accumulate". A bug fix here takes 20 to 90 requests.
  • Their stub was bigger. A 3.5 KB stub still hides most of a 16.8 KB file the agent is editing.

What to do instead: keep the prefix stableLink to this section

  • Trim big output before it goes in. The cache never sees the full text, so there's nothing to rewrite later.
  • Never edit what the model has already seen while the cache is warm.
  • Shrink once, not every turn. When a conversation must get smaller, summarize it in one step instead of removing a little at every request.

I'm not shipping ObservationPack. SoL-Pi's Action Fusion, which runs the test command in the same step as the edit, leaves the history alone, so it's the one I'll try next.