Skip to content
Round 3Empryo vs piSix models, two providers2026-08-16

Code has a shape

We compared Empryo, running with all 37 tools and its Genome (a live map of which functions call which and which files change together), against pi, a minimal agent with 4 tools. Both fixed four real bugs with merged fixes from hono and opencode, on six models from two providers, with every API call metered. On Opus 5, the most expensive model, both fixed all four bugs, and Empryo cost 27% less and finished 40% faster. Caveats: most numbers average two runs, a few are single runs, and both repos are TypeScript.

37 toolsEmpryo, full product
4 toolspi, minimal

Results at a glance

Opus 5, the priciest model: cost per round

−27%

$2.27 for Empryo against $3.10 for pi, per round. Empryo was also 40% faster, and both fixed all 4 bugs.

  • Empryo$2.27
  • pi$3.10
GPT-5.6 luna
100% fixed
pi fixed 88%. The only model where the two agents fixed a different number of bugs.
Haiku 4.5 *
−11% time
Round finished faster, same bugs fixed.
Against our previous agent (v1)
About half the cost
Opus −56%, terra −46%. luna went from fixing half the bugs to fixing all of them.
Empryo vs pi
6 models
2 providers, 4 real merged bugs, every dollar measured, August 2026.

The pricier the model, the more the map saves. Strong models use the map to go straight to the right code. Where tokens cost almost nothing, skipping the map can still come out cheaper.

* On Haiku, a helper model costing about a cent looks over the code before the main model starts. pi's one missed bug came from pi stalling, not from the test setup. pi ran in its minimal setup; Empryo ran with all its features on.

pi ran minimal. Empryo brought the map.

pi in its leanest setup. Empryo with everything turned on.

Empryo, full product

37 tools

  • live map of your code
  • memory across sessions
  • skills + helpers
  • tuned per model

pi, minimal

4 tools

  • 20-line prompt
  • no map
  • no memory
  • Opus: 27% cheaper
  • Opus: 40% faster
  • luna: more bugs fixed
  • Haiku: faster round
  • never fewer bugs fixed

Real pi setups usually add more than these four tools, and every tool adds text the model reads on every step. Empryo ran with all 37 tools and was still cheaper on Opus 5, where tokens cost the most.

Our own agent learned this the hard way

Forge, Empryo's main agent, version 1 against version 2: same bugs, same models, same keys. A shorter cost stroke means cheaper.

Opus 5−56%

cost per round

  • v1$5.12
  • v2$2.27

GPT-5.6 terra−46%

cost per round

  • v1$1.58
  • v2$0.847

GPT-5.6 luna−10%

cost per round

  • v1$0.110
  • v2$0.099

GPT-5.6 luna50% to 100%

bugs fixed

  • v150%
  • v2100%

Haiku 4.575% to 88%

bugs fixed

  • v175%
  • v288%

Forge v1 showed every model the whole code map on every step, and paid for it. v2 adapts to the model: strong models get a short summary, small models get more detail. It also limits how much text one tool result can add to the conversation. The cost per round dropped by about half.

On the most expensive model, the map pays off

Opus 5, cost per task. All three fixed every bug, so the costs compare like for like. A shorter stroke means cheaper.

Whole round

4 bugs in 2 repos

  • Empryo v1$5.12
  • Empryo v2$2.27
  • pi$3.10

Cookie mangling

hono, easy

  • Empryo v1$0.945
  • Empryo v2$0.468
  • pi$0.269

TrieRouter regexp

hono, hard

  • Empryo v1$0.845
  • Empryo v2$0.570
  • pi$0.634

Symlinked grep root

opencode, big repo

  • Empryo v1$1.93
  • Empryo v2$0.762
  • pi$1.19

Revert by ID

opencode, hard

  • Empryo v1$1.41
  • Empryo v2$0.472
  • pi$1.01

An early v2 build lost on Opus. One search result the size of a small book stayed in the conversation and was re-read on every later step. v2 now caps how long a tool result can be. With that change, Empryo beat pi by 27% on cost and 40% on time over the round, using 60 requests to pi's 80. pi was still cheaper on the easy cookie bug.

The right file

4.3×

faster to the right file

In the big repo, two files share a name and only one holds the bug. An agent that searches text has to guess which. Empryo followed the map and found it in 86 seconds. pi searched for 369. Our v1 agent picked the wrong file and patched it.

  • Empryo86s
  • pi369s

Code is more than text.

Functions call other functions. Files change together. Give your agent that map and each dollar goes further.

The fine print

What the numbers mean

Plain words for every measure on this page.

A run
One agent, one model, one bug, one attempt. A cell in the tables below is one run.
Fixed
The bug only counts as fixed when tests the agent never saw pass on its code. We add those tests after the agent has finished, so it cannot aim at them.
Step
One round trip to the model: the agent sends the conversation so far and gets back its next move. Each step sends the whole conversation again, so every step you avoid is one you don't pay for.
Tokens
The pieces of text a model reads and writes, a word or part of a word each. Providers charge per token.
Metered
Every request to the model went through a small recording proxy that counted the tokens and dollars on the wire, as a third record next to the agent's report and the console.
Time
Real time from start to finish, as on a stopwatch. Also called wall clock.
Genome
Empryo's live map of your code: which functions call which, and which files change together. The agent reads the map instead of guessing from text search.

Don't trust our numbers. Run them.

The tasks, hidden tests and metering for this round are all in the public bench repo, so you can rerun every model yourself.