Skip to content

Code has a shape. The bill knows.

One search result cost us $1.14, more than a rival agent spent on an entire bug fix. So we rebuilt Empryo's engine, tested it against the most minimal agent we know, and metered every dollar across six models. Pictures inside.

Code has a shape. The bill knows.

One search result cost us $1.14.

Not the reasoning, and not the fix. One search that dumped 86,000 tokens into the conversation. The model needed two lines of it. But every later step re-sent the whole conversation to the model, dump included, because that is how agent sessions work. You pay for tool output once when it arrives, and again on every step until the session ends.

From the request log: tool output doesn't disappear after it's read. It gets re-sent, and re-billed, on every later step.
From the request log: tool output doesn't disappear after it's read. It gets re-sent, and re-billed, on every later step.

That receipt is why we rebuilt Forge, the engine that decides what context and tools Empryo sends to the model on each step. Then we tested the rebuild under strict conditions.

Every number below comes from recorded runs: Round 1, Round 2, and the new Round 3 results.

The testLink to this section

We used four bugs that were really reported and really fixed in hono and opencode. The tests from each fix's pull request were hidden from both agents. Git history was removed so the answer couldn't be looked up. Six models, two providers, 96 metered runs.

pi in its minimal setup, Empryo in its default setup: what each agent sends the model before any work starts.
pi in its minimal setup, Empryo in its default setup: what each agent sends the model before any work starts.

Partway through, we found that pi's cost meter was wrong for GPT models, so we upgraded pi to get correct numbers. Making your opponent's bill more accurate is the cheapest credibility there is.

What the bill taught usLink to this section

Forge v1, the old engine, gave every model the same context: the whole code map and a full description of every tool, on every step. The request logs made it clear that most of this was being paid for and not used.

Forge v2 in one picture: strong models get a short overview, small models get more guidance and a helper to scout the code, and every tool result has a size budget.
Forge v2 in one picture: strong models get a short overview, small models get more guidance and a helper to scout the code, and every tool result has a size budget.

Same bugs, same models, same keys. Switching engines alone cut cost by 56% on Opus and 46% on GPT-5.6 terra. GPT-5.6 luna went from fixing half the bugs to fixing all of them.

One thing v2 deliberately leaves out: hints tuned to the benchmark. We wrote some, and they worked. We deleted them, because an engine tuned to a test is a worse engine with a better score.

The run that explains itLink to this section

Two files with the same name, one bug. A text search has to guess which file matters. The code structure tells you.
Two files with the same name, one bug. A text search has to guess which file matters. The code structure tells you.

That one run shows the whole idea. Code isn't just a wall of text that happens to compile. Functions call other functions, and some files always change together. An agent that can see those links doesn't have to guess.

The resultsLink to this section

Round 3 wins. The full results, including every loss, are on the results page.
Round 3 wins. The full results, including every loss, are on the results page.

The parts that didn't go our way, because a benchmark without them is an ad. On cheap and mid-priced models, pi's tiny prompt keeps the cost advantage. When a million tokens costs a quarter, a twenty-line prompt is hard to beat. pi's totals are also very consistent between runs, while ours vary (on Sonnet, $1.20 in one run and $1.47 in another). Making our costs more consistent is the next goal.

But the pattern supports the idea behind the rebuild: the stronger the model, the more a code map pays off. Frontier models that are given a map take fewer, better steps, and at frontier prices fewer steps saves real money. Small models do better when the map is handed to them in smaller pieces. How you deliver code knowledge to a model matters as much as having it.

Want to check our work? The tasks, hidden tests, and metering are all in the public benchmark repo. Current Empryo builds run Forge v2; in the desktop app, the feather icon next to the prompt box shows it is active. Either way, the bill is the judge.