I tested NVIDIA's SoL-Pi pruning on real bugs. It cost 30% more.
On real Hono bugs, SoL-Pi's ObservationPack cost 30% more. No pruning rule I tried beat one habit: never change what the model has already seen.
15 articles, newest first.
On real Hono bugs, SoL-Pi's ObservationPack cost 30% more. No pruning rule I tried beat one habit: never change what the model has already seen.
Jev is a decision model from TypeSafe AI. It answers typed questions with probabilities instead of text, in a few hundred milliseconds for a fraction of a cent. I wired it into five places in Empryo, measured it against frontier models, and kept it only where it won.
A practical trial for Claude Code, Codex, Cursor, OpenCode, and Empryo. Compare accepted fixes, full costs, review time, and failed attempts.
Follow a symbol rename through an AI coding harness. See where context, tools, permissions, tests, and session state affect the result.
A Morph is one JSON file that adds a panel to the Empryo desktop app. A notebook, your tracker as a drag-and-drop board, a timer that chimes. The app checks the file and draws it with its own components. You use the panel directly; the agent only builds it.
One QA system, two ways to run it. I run it as a fixed suite and get the same answer every time. An agent runs it to explore, picks targets from the code map, and writes the check it wishes had existed. 105 defects claimed, 84 survived two skeptical reviewers.
One search result cost us $1.14, more than a rival agent spent on an entire bug fix. So we rebuilt Empryo's engine, tested it against the most minimal agent we know, and metered every dollar across six models. Pictures inside.
Empryo's agent writes its own notes and ties each one to the files it's about. Open a file and its notes come back, labeled with why. Rename the file and they fade. Pin what matters; notes nobody uses archive themselves. Memory that follows your repo instead of sitting in the prompt.
My first prompt pre-pass guessed at causes and dropped the fix rate from 3/3 to 1/3. The version that only points at files scored 3/3 at 11% less spend. Same model, same bug. The difference was removing every place a guess could go.
A formatter in a side app ranked as the most important symbol in my whole codebase. Static analysis said it mattered; my git history said it didn't. When I let the map learn from how I actually work, co-change pairs went from 86 to 3,239 and the formatter fell to #103.
Ten coding agents ranked on what they fix, what they cost, how they understand your code, and how much freedom you keep. Where I have head-to-head data, the costs come from real provider bills. Disclosure, I build #2.
pi said $6.25. Anthropic billed $9.19. Empryo said $7.06 and was billed $7.08. An agent's own cost figure is only a claim until you check it against the provider's bill. Here is how I check.
We stopped writing benchmark bugs and used real merged fixes from hono, zod and ky, all merged recently enough that models are unlikely to have seen them. 7/10 vs 6/10, one 8-minute hang, and one bug that beat everyone.
A failed edit costs three tool calls: the miss, the re-read, the retry. Empryo's ast_edit targets a function or type by name and changes the syntax tree directly. No search string, no line numbers, nothing to mismatch. It started as my Master's thesis.
Before your first message, Empryo's agent already knows every exported symbol, who calls it, and how many files depend on it. The Genome is that map. It is ranked, included in the agent's context, and updated after every edit.