Skip to content

I gave my test suite to the agents

One QA system, two ways to run it. I run it as a fixed suite and get the same answer every time. An agent runs it to explore, picks targets from the code map, and writes the check it wishes had existed. 105 defects claimed, 84 survived two skeptical reviewers.

I gave my test suite to the agents

Your tests are green and a user still finds a bug in five minutes. Two of those got past me, and no test I would have written could have caught either one.

The first was an installer. I ran it on a fresh Debian machine with no Empryo, no Bun, and no TERM environment variable, because nothing sets one when you open a shell inside a container.

Downloaded. Checksummed. Signature verified against the pinned RSA-4096 key. Then dead, with nothing installed.
Downloaded. Checksummed. Signature verified against the pinned RSA-4096 key. Then dead, with nothing installed.

The script did everything right. Then it ran clear to tidy the screen before printing its welcome message. clear needs TERM. It exited with an error, set -e stopped the script, and the verified download was never unpacked.

The second was worse, because nothing failed at all. The desktop app shipped with no right-click menu in any text field. You couldn't paste with the mouse. Every test stayed green, because no test checked for it. A user found it immediately.

Those are two different failures, and they need two different fixes. The first got through because my tests import my code, so they never run the installer. The second got through because a test can only repeat a question someone already thought to ask.

This post describes the QA system I built in response. Empryo calls it the Immune system.

One system, two ways to run itLink to this section

The same seeds, the same real machines, the same two reviewers. What changes is who picks the targets.
The same seeds, the same real machines, the same two reviewers. What changes is who picks the targets.

I built one system that can be driven two ways.

I run it as a fixed suite. It runs 260 checks across three operating systems and the three ways you can use Empryo (desktop, terminal UI and headless), and gives the same answer every time. This half repeats known questions. It's the half I trust.

An agent runs it to explore. The machines and reviewers are the same, but the agent chooses what to test. It uses the product the way a person would and writes the check it wishes had existed. This half finds new problems. It's the half I double-check.

Neither side can add its own findings to the suite. That rule is what makes the rest usable.

What the fixed suite runsLink to this section

Each check runs on one platform and one way of using Empryo. A general sweep inspects every result for crashes and garbled output, whatever the check was about.
Each check runs on one platform and one way of using Empryo. A general sweep inspects every result for crashes and garbled output, whatever the check was about.

Each check (I call them seeds) states what must be true: what to exercise, on which operating systems, in which of the three ways to use Empryo, and one sentence naming the kind of bug it guards against. A run executes every seed on every platform it applies to, in parallel.

Only the model is faked. A scripted stand-in answers instead of a real provider, so the agent loop, the tool calls and the network traffic all run for real, offline and for free. Everything else is real: a real terminal, a real desktop window, a real Linux machine, and a real Windows x64 machine.

The Windows machine has to be real x64 hardware. An ARM64 Windows virtual machine can't load x64 native modules, so a pass there would be a pass for a product nobody ships. It's only used for quick smoke tests.

The part I didn't expect to matter most is a general sweep that inspects every result, whatever the check was looking for. It flags stack traces, terminal escape codes printed as text, locked-database errors, and processes that had to be killed. Most surprises show up there. A check about resuming an unknown session led me to a crash in the runtime's HTTP client that had nothing to do with sessions and was killing about 40% of runs.

How an agent picks what to testLink to this section

Empryo's own code map turned out to be a QA tool.

Grep ranks by what you already thought to type. The graph ranks by how many callers reach a function and how far a defect there would spread.
Grep ranks by what you already thought to type. The graph ranks by how many callers reach a function and how far a defect there would spread.

An agent has to choose a target. With grep, it picks whatever it can spell, usually the file whose name sounds most like the feature. That file is rarely the one that matters.

The Genome, Empryo's live map of your code, answers a better question. Its tools show which files depend on a given file, which files usually change with it, and how far a change there would spread. They find exports nothing uses. And they list every caller that can reach a given function, with file and line.

That last one does two jobs, which I only noticed after using it. The callers that reach a function also tell you which platforms and ways of using Empryo a bug in it can affect. That's exactly what a new check has to declare. The tool that finds the bug also fills in half of the check.

One rule about the agents matters more than any other: none of them can start another agent. Every handoff means the agent writes down what it found and a person starts the next step. That rule stays. An agent that could add its own findings straight to the suite would fill it with checks that fail on correct behaviour.

Nothing becomes a check until it survives reviewLink to this section

Two reviewers per claim, both hostile, both defaulting to refuted. One argues the behaviour is documented intent. The other reproduces by hand on the shipped build.
Two reviewers per claim, both hostile, both defaulting to refuted. One argues the behaviour is documented intent. The other reproduces by hand on the shipped build.

Agents that hunt for bugs want to find bugs. Left alone, they produce checks that fail on correct behaviour. Soon a failing run means nothing.

So every claim goes past two reviewers before it becomes a check. One asks whether the behaviour is actually what the product is documented to do. The other ignores the test setup and reproduces the problem by hand on the released build, in a throwaway folder, the way a user would. Both start from the assumption that the claim is wrong.

One run came back with 23 failing checks on Windows. Five were real. Sixteen came from one outdated binary the test machine was stuck on, and two from checks written with Unix-only assumptions. Throwing out those eighteen false alarms is the most valuable thing the review does. Rejecting a claim is a good outcome here.

A QA report where nothing is ever rejected is measuring its reviewers, not its findings. So the record only grows, and it lists rejections next to confirmations: 105 claimed, 84 confirmed, 19 rejected, 2 disputed.

What it caughtLink to this section

Three defects, all fixed now, each leaving a permanent check behind:

what a user was losing
edit_file wrote through a read-only fileYou chmod 444 a generated file. The agent edited it anyway, was told the edit succeeded, and kept reasoning about a file it should never have touched.
Editing a file lost its permissionsThe edit replaced the file with a new one, so its permissions reset to the default. Every edited script quietly stopped being executable.
--resume on an unreadable session answered with no historyThe session file couldn't be read, so the agent started empty, answered confidently, and billed you for a turn that had forgotten the conversation.

None of these threw an error. None exited with a failure code. All three looked like success from the outside, which is why unit tests missed them.

The first-day checks, which install Empryo on a clean machine and watch what a new user sees, found four more. There was the clear failure above. There was a tar: warning printed once per file, because macOS had packed extra file attributes into the archive. There was a warning that Empryo needs Bun installed, when it ships its own. And curl's progress bar was spilling into piped logs.

One more: a release build failed its own startup check. The binary was fine. On startup it cleaned up leftover language servers before answering --version. On a machine with old data, that cleanup triggered a migration that took longer than the check allowed. The fix was moving one call.

What isn't coveredLink to this section

373 capabilities the product exposes. 96 have a check aimed at them. The gap is the honest number, and it is the one that moves.
373 capabilities the product exposes. 96 have a check aimed at them. The gap is the honest number, and it is the one that moves.

Empryo has 373 capabilities you can reach. 96 have a check. That leaves 277 that a run would pass straight over. The coverage list is generated from the product itself, so I can't quietly stop counting the gaps.

Think of the eight stages of using a tool, from installing it to uninstalling it. Three are tested end to end on a machine that has never seen Empryo. The other five aren't. A stage only counts if the test starts from a clean machine. Setting things up first and then calling it covered is the easiest lie to tell yourself, so the system doesn't accept it.

What I haven't builtLink to this section

I want to be precise here, because this is the part people skip to.

This system doesn't fix itself. Agents find defects and write the checks that guard against them, and they're good at it. But every run is a command someone types. Every finding is passed to a reviewer by hand. Every fix is still a person deciding what the right behaviour is. Nothing runs unattended. Calling it self-healing would be exactly the kind of claim this system exists to catch.

What's true is narrower: the loop closes. A defect is found, challenged by a reviewer that defaults to no, fixed, and then guarded by a check written from that exact defect. Every round leaves the suite bigger and the gap smaller. 84 confirmed defects across three rounds is a good start, and also a slightly alarming statement about what was in there.

What would it take to go further? Scheduled runs, so testing isn't something I have to remember. Agents that hand off to each other under clear rules instead of through me. And a fix step that proposes a patch and proves it against the check that failed, then stops, because the final call should stay with a person.

Try the idea on your own appLink to this section

None of this is specific to Empryo. It's specific to having a product and wanting to know whether it works.

The ideas are general. Test the real thing you ship, not the code you import. Use a map of your code to choose what to test instead of guessing. Don't let a claim become a test until a skeptical reviewer fails to knock it down. Publish what isn't covered next to what is. The Empryo-specific part is that the code map is already built in, so the tool doing the checking can read the code it's checking.

The current numbers, including the gaps, are at empryo.com/immunity and update on every deploy.