Skip to content

The reviewer

A separate agent with a clean context reads what changed and returns PASS, FAIL or PARTIAL. Run it on demand, loop it until it passes, or use it in CI.

How it works5 minutes to read
Copy & share
On this page7 sections

An agent that just spent forty steps writing a patch is a poor judge of it. It remembers why each line seemed right when it wrote it.

The reviewer is a separate agent that starts with none of that history. It opens the files that changed, reads them fresh, and ends with a one-line verdict your tools can act on:

VERDICT: PASS — [what was verified]
VERDICT: FAIL — [file:line, what is wrong]
VERDICT: PARTIAL — [what could not be verified and why]

You can use it three ways: /review runs it once, `/goal` runs it after every round, and --review turns the verdict into an exit code for CI.

Run a reviewLink to this section

/review
/review check the error paths, not the happy path

The text after /review is optional. If you add it, the reviewer checks that first. If you leave it out, it checks the changed files for correctness.

A review can't start while work is still in progress: while a reply is streaming, a background agent is still writing files, or a goal loop is running. Reviewing code that keeps changing would give you a verdict about code that no longer exists. The review button tells you what's blocking it and becomes available again when the work finishes.

You can type while a review runs. Your message reaches the reviewer at its next step. These notes are kept out of the conversation, so the coder doesn't later mistake them for instructions meant for it.

What gets reviewedLink to this section

Empryo checks three places in order and reviews the first one that has changes:

OrderSourceWhat the reviewer sees
1This tab's editsOnly files this tab changed
2This session's editsEvery file any tab changed this session
3gitEverything uncommitted, including new files (.gitignore still applies)

With four tabs open, each tab's /review checks its own work. If nothing changed, no model is called and it costs nothing.

A review covers at most 30 files. The verdict says which of the three sources it used, so you can tell "the four files this tab wrote" from "everything uncommitted".

For the first two sources, each file shows +N/−M compared with how it looked before this session first changed it. Those snapshots are kept in memory only. After a restart, a restored tab falls back to git, and the report says so.

Deleted files are included and marked as deleted. The reviewer doesn't try to open them. Instead it checks what used to import or call them, because a deletion that leaves callers behind is exactly the kind of bug you want caught.

Acting on a verdictLink to this section

AppWhere you read itHow you act on it
DesktopThe reviewer bar above the message box. Expand it to see past verdictsapply sends the findings to the agent. re-review runs a new review on the same files
Terminal UIThe banner above the message box, and the report in the conversation/review apply, /review list

After a single /review, nothing changes until you choose to apply the findings.

Each tab keeps its last 8 verdicts and their reports, including after a reload or session restore. A re-review is given the previous report as something to check, not as fact, so the new reviewer has to confirm every claim again.

Review until it passesLink to this section

Doing it by hand means: review, read, apply, wait, review again. /review pass does that cycle for you. Every verdict that isn't PASS goes to the coder as a normal turn, and as soon as that turn finishes, a new reviewer runs.

/review pass
/review pass the error paths, not the happy path
/review stop

On desktop, the review button opens a menu with Review once and Review till pass. Both accept the same optional instructions from the message box. The reviewer bar shows the current round (for example round 2/5) and a stop loop button. In the terminal UI, a finished verdict shows a ↻ till pass option next to ↯ apply.

It stops on PASS, at the round limit (5 by default, never more than 12), on a review error, or when you stop it. Ctrl+X or the stop button ends it at any point, including the reviewer that's currently running. Typing while it runs reaches whichever agent is active: the reviewer or the coder.

Each round is a real coder turn plus a real review, and both count toward the tab's costs. A loop that reaches its limit runs five reviews and four coder turns, so give it clear instructions.

How it differs from `/goal`: /goal works toward a goal you state up front and switches to stronger models when it stalls. /review pass checks code that already exists, with no goal other than surviving a fresh review.

In CILink to this section

bash
empryo --headless --review
empryo --headless --review "check the migration for data loss"
empryo --headless --review --json
Exit codeMeaning
0PASS, or nothing to review (no changes, no model called)
3FAIL
4PARTIAL
1 / 2 / 130Error / timeout / aborted
yaml
- name: Review the diff
  run: empryo --headless --review --quiet

With no changes, the JSON output has verdict: null, never PASS. Nothing was reviewed, and reporting a pass would tell your pipeline the code had been checked.

Which model reviewsLink to this section

The reviewer uses the review slot in the task router. If that's empty, it uses goalReview, then the model your tab is already using. This is a good place to spend on a strong model: a reviewer that misses a bug costs more than it saves. Review costs appear in the tab's totals, so you can see what each verdict cost (cost tracking).

Typecheck and testsLink to this section

/review doesn't run your typecheck or tests, so a manual review finishes in seconds. The goal loop does run them between rounds, where a slower, stricter check is worth the time.