Skip to content
15 articles since May 2025

Blog

How Empryo works, why it is built the way it is, and what testing it taught us.

Every article

Filter articles by topic

15 articles, newest first.

Engineering

9 min read

Software needs decisions, not chat: testing TypeSafe's Jev inside Empryo

Jev is a decision model from TypeSafe AI. It answers typed questions with probabilities instead of text, in a few hundred milliseconds for a fraction of a cent. I wired it into five places in Empryo, measured it against frontier models, and kept it only where it won.

Engineering

8 min read

I stopped building panels and let you morph the app

A Morph is one JSON file that adds a panel to the Empryo desktop app. A notebook, your tracker as a drag-and-drop board, a timer that chimes. The app checks the file and draws it with its own components. You use the panel directly; the agent only builds it.

Engineering

9 min read

I gave my test suite to the agents

One QA system, two ways to run it. I run it as a fixed suite and get the same answer every time. An agent runs it to explore, picks targets from the code map, and writes the check it wishes had existed. 105 defects claimed, 84 survived two skeptical reviewers.

Engineering

3 min read

Code has a shape. The bill knows.

One search result cost us $1.14, more than a rival agent spent on an entire bug fix. So we rebuilt Empryo's engine, tested it against the most minimal agent we know, and metered every dollar across six models. Pictures inside.

Architecture

10 min read

Your agent's memory is a search engine. Code memory should be a neighbor.

Empryo's agent writes its own notes and ties each one to the files it's about. Open a file and its notes come back, labeled with why. Rename the file and they fade. Pin what matters; notes nobody uses archive themselves. Memory that follows your repo instead of sitting in the prompt.

Architecture

5 min read

My code map was wrong about my own repo. So I taught it to learn.

A formatter in a side app ranked as the most important symbol in my whole codebase. Static analysis said it mattered; my git history said it didn't. When I let the map learn from how I actually work, co-change pairs went from 86 to 3,239 and the formatter fell to #103.

Benchmarks

4 min read

I benchmarked two AI agents. Then I read the bill.

pi said $6.25. Anthropic billed $9.19. Empryo said $7.06 and was billed $7.08. An agent's own cost figure is only a claim until you check it against the provider's bill. Here is how I check.

Benchmarks

5 min read

Five real bugs, two agents, and the one nobody could fix

We stopped writing benchmark bugs and used real merged fixes from hono, zod and ky, all merged recently enough that models are unlikely to have seen them. 7/10 vs 6/10, one 8-minute hang, and one bug that beat everyone.

Engineering

4 min read

Why an AI agent should edit symbols, not strings

A failed edit costs three tool calls: the miss, the re-read, the retry. Empryo's ast_edit targets a function or type by name and changes the syntax tree directly. No search string, no line numbers, nothing to mismatch. It started as my Master's thesis.

Architecture

3 min read

The Genome - a live structural map of your codebase

Before your first message, Empryo's agent already knows every exported symbol, who calls it, and how many files depend on it. The Genome is that map. It is ranked, included in the agent's context, and updated after every edit.