Code has a shape. The bill knows.

One search result cost us $1.14.
Not the reasoning, and not the fix. One search that dumped 86,000 tokens into the conversation. The model needed two lines of it. But every later step re-sent the whole conversation to the model, dump included, because that is how agent sessions work. You pay for tool output once when it arrives, and again on every step until the session ends.
That receipt is why we rebuilt Forge, the engine that decides what context and tools Empryo sends to the model on each step. Then we tested the rebuild under strict conditions.
Every number below comes from recorded runs: Round 1, Round 2, and the new Round 3 results.
The testLink to this section
We used four bugs that were really reported and really fixed in hono and opencode. The tests from each fix's pull request were hidden from both agents. Git history was removed so the answer couldn't be looked up. Six models, two providers, 96 metered runs.
Partway through, we found that pi's cost meter was wrong for GPT models, so we upgraded pi to get correct numbers. Making your opponent's bill more accurate is the cheapest credibility there is.
What the bill taught usLink to this section
Forge v1, the old engine, gave every model the same context: the whole code map and a full description of every tool, on every step. The request logs made it clear that most of this was being paid for and not used.
Same bugs, same models, same keys. Switching engines alone cut cost by 56% on Opus and 46% on GPT-5.6 terra. GPT-5.6 luna went from fixing half the bugs to fixing all of them.
One thing v2 deliberately leaves out: hints tuned to the benchmark. We wrote some, and they worked. We deleted them, because an engine tuned to a test is a worse engine with a better score.
The run that explains itLink to this section
That one run shows the whole idea. Code isn't just a wall of text that happens to compile. Functions call other functions, and some files always change together. An agent that can see those links doesn't have to guess.
The resultsLink to this section
The parts that didn't go our way, because a benchmark without them is an ad. On cheap and mid-priced models, pi's tiny prompt keeps the cost advantage. When a million tokens costs a quarter, a twenty-line prompt is hard to beat. pi's totals are also very consistent between runs, while ours vary (on Sonnet, $1.20 in one run and $1.47 in another). Making our costs more consistent is the next goal.
But the pattern supports the idea behind the rebuild: the stronger the model, the more a code map pays off. Frontier models that are given a map take fewer, better steps, and at frontier prices fewer steps saves real money. Small models do better when the map is handed to them in smaller pieces. How you deliver code knowledge to a model matters as much as having it.
Want to check our work? The tasks, hidden tests, and metering are all in the public benchmark repo. Current Empryo builds run Forge v2; in the desktop app, the feather icon next to the prompt box shows it is active. Either way, the bill is the judge.