Code has a shape
We compared Empryo, running with all 37 tools and its Genome (a live map of which functions call which and which files change together), against pi, a minimal agent with 4 tools. Both fixed four real bugs with merged fixes from hono and opencode, on six models from two providers, with every API call metered. On Opus 5, the most expensive model, both fixed all four bugs, and Empryo cost 27% less and finished 40% faster. Caveats: most numbers average two runs, a few are single runs, and both repos are TypeScript.
Results at a glance
Opus 5, the priciest model: cost per round
−27%$2.27 for Empryo against $3.10 for pi, per round. Empryo was also 40% faster, and both fixed all 4 bugs.
- GPT-5.6 luna
- 100% fixed
- pi fixed 88%. The only model where the two agents fixed a different number of bugs.
- Haiku 4.5 *
- −11% time
- Round finished faster, same bugs fixed.
- Against our previous agent (v1)
- About half the cost
- Opus −56%, terra −46%. luna went from fixing half the bugs to fixing all of them.
- Empryo vs pi
- 6 models
- 2 providers, 4 real merged bugs, every dollar measured, August 2026.
The pricier the model, the more the map saves. Strong models use the map to go straight to the right code. Where tokens cost almost nothing, skipping the map can still come out cheaper.
* On Haiku, a helper model costing about a cent looks over the code before the main model starts. pi's one missed bug came from pi stalling, not from the test setup. pi ran in its minimal setup; Empryo ran with all its features on.
pi ran minimal. Empryo brought the map.
pi in its leanest setup. Empryo with everything turned on.
Empryo, full product
37 tools
pi, minimal
4 tools
- Opus: 27% cheaper
- Opus: 40% faster
- luna: more bugs fixed
- Haiku: faster round
- never fewer bugs fixed
Real pi setups usually add more than these four tools, and every tool adds text the model reads on every step. Empryo ran with all 37 tools and was still cheaper on Opus 5, where tokens cost the most.
Our own agent learned this the hard way
Forge, Empryo's main agent, version 1 against version 2: same bugs, same models, same keys. A shorter cost stroke means cheaper.
Opus 5−56%
cost per round
GPT-5.6 terra−46%
cost per round
GPT-5.6 luna−10%
cost per round
GPT-5.6 luna50% to 100%
bugs fixed
Haiku 4.575% to 88%
bugs fixed
Forge v1 showed every model the whole code map on every step, and paid for it. v2 adapts to the model: strong models get a short summary, small models get more detail. It also limits how much text one tool result can add to the conversation. The cost per round dropped by about half.
On the most expensive model, the map pays off
Opus 5, cost per task. All three fixed every bug, so the costs compare like for like. A shorter stroke means cheaper.
Whole round
4 bugs in 2 repos
Cookie mangling
hono, easy
TrieRouter regexp
hono, hard
Symlinked grep root
opencode, big repo
Revert by ID
opencode, hard
An early v2 build lost on Opus. One search result the size of a small book stayed in the conversation and was re-read on every later step. v2 now caps how long a tool result can be. With that change, Empryo beat pi by 27% on cost and 40% on time over the round, using 60 requests to pi's 80. pi was still cheaper on the easy cookie bug.
The right file
faster to the right file
In the big repo, two files share a name and only one holds the bug. An agent that searches text has to guess which. Empryo followed the map and found it in 86 seconds. pi searched for 369. Our v1 agent picked the wrong file and patched it.
Code is more than text.
Functions call other functions. Files change together. Give your agent that map and each dollar goes further.
The fine print
What the numbers mean
Plain words for every measure on this page.
- A run
- One agent, one model, one bug, one attempt. A cell in the tables below is one run.
- Fixed
- The bug only counts as fixed when tests the agent never saw pass on its code. We add those tests after the agent has finished, so it cannot aim at them.
- Step
- One round trip to the model: the agent sends the conversation so far and gets back its next move. Each step sends the whole conversation again, so every step you avoid is one you don't pay for.
- Tokens
- The pieces of text a model reads and writes, a word or part of a word each. Providers charge per token.
- Metered
- Every request to the model went through a small recording proxy that counted the tokens and dollars on the wire, as a third record next to the agent's report and the console.
- Time
- Real time from start to finish, as on a stopwatch. Also called wall clock.
- Genome
- Empryo's live map of your code: which functions call which, and which files change together. The agent reads the map instead of guessing from text search.