Cost tracking
Live USD cost tracking per model and per sub-agent, with built-in pricing for every major provider, cache-aware billing, and a /context breakdown dashboard.
Copy & share
Loading sections…
Connect MCP or install the Empryo skillSection exports contain only that heading’s content. Markdown and text links fetch the selected content directly, without the rest of the page.
Empryo tracks every token spent and prices it in real time. The status bar shows the running total in USD. /context opens a dashboard with the per-model breakdown.
What gets tracked
- Prompt tokens (uncached input).
- Completion tokens (model output).
- Cache-write tokens (billed at a higher rate by most providers).
- Cache-read tokens (billed at a discount).
- Subagent tokens tracked separately from the main agent.
- Per-model breakdown when the task router mixes providers.
Providers with built-in pricing
Pricing tables ship for the major providers, updated against their public price lists:
| Provider | Notes |
|---|---|
| Anthropic | Claude Opus/Sonnet/Haiku with cache-write and cache-read rates |
| OpenAI | GPT-5.4, GPT-4.1, o3, o4-mini |
| Gemini 2.5 Pro/Flash, Gemini 3 Flash/Pro | |
| DeepSeek | V4 (pro/flash), V3, R1; chat + reasoner aliases |
| Groq | Llama 3.3, Llama 4 Scout, Qwen3, GPT-OSS |
| Mistral | Mistral Large/Medium/Small, Codestral, Magistral, Ministral, Pixtral, Devstral |
| Fireworks | Tier-based pricing (Mixtral, Llama 70B+, DeepSeek) |
| GitHub Copilot | Premium-request multiplier-based estimation |
| OpenRouter | Live pricing from the catalog |
| GitHub Models | Per-token via multipliers |
| Ollama, LM Studio, OpenCode free models | $0.00 |
Custom providers default to a conservative estimate. Unknown models fall back to Sonnet-tier pricing as a safety floor.
Why it matters
Two tactics cut cost dramatically:
- Mix models. Haiku for spark agents, Sonnet for ember agents, Flash for compaction. A task that would cost $0.25 on Sonnet often runs for $0.05 when the exploration phase routes through Haiku.
- Use caching. Cache reads are 10x cheaper on Anthropic, up to 50% off on Groq/Fireworks. Empryo structures the system prompt and the Genome for maximum cache hits - typical cache-hit rates exceed 60%.
UI
The status bar shows the running total in USD. Compact mode shows tokens plus a dollar figure. /context opens the detailed view: per-model usage, cache ratio, subagent spend, and the compaction history.
/usage opens the spend ledger — today, this week, this month, per model — plus a cumulative prompt-cache hit rate so you can see caching working across sessions, not just in the current turn. A healthy rate means the cached prefix is holding; a low one means something keeps re-writing it.
Use /router to assign cheap models to cheap tasks.
Flat-rate plans are measured differently
Everything above prices tokens in dollars. Two routes do not bill per token at all — they spend an allowance, and the only limit that matters is how much of it is left:
- **
proxy/*** — a rolling subscription window through the proxy add-on. - **
copilot/*** — a monthly Copilot allowance on your GitHub seat.
Those meters sit alongside the dollar figures rather than replacing them, because one session can do both: a gateway lane billing credits and a relay lane burning a weekly cap.