EmpryoEmpryo.beta
console-audited · separate keys · reproducible

Every agent says it's fast and cheap. The bill decides.

Real bugs, reported the way humans report them. Each agent runs on its own Anthropic API key, graded by hidden acceptance tests — then its self-reported cost is audited against the provider's actual bill. Both numbers get published.

latest · round 03 · opus 5 · aug 2026
cost / round−27%
empryo$2.27
pi$3.10
wall clock−40%
empryo5m 45s
pi9m 35s
shorter bar = cheaper & faster · both fixed 4/4 · full board →
the audit · 19 bugs, both agents · 2 rounds · anthropic consoleconsole-verified
empryosaid $8.19billed $8.21✓ matched — both audits
pisaid $7.67billed $10.77✗ missed 29% of its bill
fixes 15/19 vs 13/19billed cost −24%wall clock −38%
round 03 · vs pi · six models · 2026-08-16new

Code has a shape: six models, two providers

−27%
Opus cost
−40%
Opus wall clock
−56%
vs Forge v1

The bet: a 37-tool product with a live map of your code against the leanest agent alive, benched at its cheapest possible shape. Four real merged bugs, models from Haiku to Opus 5, every dollar measured. The smarter the model, the cheaper Empryo gets.

protocol traceforge-v2
"do agents even need tools?" — minimalism, tested at its floor
·hidden acceptance · 4 real merged bugs · 6 models, 2 providers · 96 metered runs
Opus: −27% cost, −40% wall · luna: 100% vs 88% fixed · accuracy never behind
six modelstwo providers4 merged bugsv2 −56% vs our own v1luna 100% vs 88%Haiku −11% wall clock
see full results →
round 02 · vs pi · real world · 2026-07-16

Five real bugs from hono, zod and ky

−23%
billed cost
−28%
steps
−32%
wall clock

No more fixtures. Real merged fix PRs, the original issue text verbatim, git history scrubbed so the answer is unreachable, every fix past the models' training cutoffs. Graded by the fix PR's own regression tests.

protocol tracereal-world
"RPC headers not merging when function" — hono #5089, verbatim
·hidden acceptance · the fix PR's own regression tests, dropped in after the run
7/10 vs 6/10 fixed · billed $7.08 vs $9.19 · wall −32%
post-cutoff fixeshistory-scrubbed reposhaiku sweep 3/5 vs 2/5only empryo cracked zod first tryky bug unsolved 0/4crash-invisible spend caught
see full results →
round 01 · vs pi · hookboard · 2026-07-11

Same bugs, same models, and the bill disagrees

5.7× fewer
input tokens
−28%
billed cost
−57%
wall clock

Three committed bugs in a realistic TypeScript webhook app, reported the way humans report them. One hid a trap: a stale comment claiming ids are ULIDs. Grep believes comments; the graph knows better.

protocol tracehookboard
"the order is all over the place — definitely not newest first"
·trap: stale ULID comment · pi burned 35 steps believing it, haiku+genome took 10
8/9 vs 7/9 fixed · 5.7× fewer input tokens billed · cost −28%
3 planted bugshidden checksseparate keysself-report auditpi undercounted 73%empryo exact to the cent
see full results →
04vs opencodeNext opponent — queued

Don't trust our numbers. Run them.

Two API keys and one command reproduce any round — harness, fixtures, hidden checks and raw results are public. Task validation is free.

bun harness-real/run-real.ts --tiers haiku,opus