concepts

Cost tracking

Live USD cost tracking per model and per sub-agent, with built-in pricing for every major provider, cache-aware billing, and a /context breakdown dashboard.

Copy & share

Loading sections…

Connect MCP or install the Empryo skill

Section exports contain only that heading’s content. Markdown and text links fetch the selected content directly, without the rest of the page.

Empryo tracks every token spent and prices it in real time. The status bar shows the running total in USD. /context opens a dashboard with the per-model breakdown.

What gets tracked

  • Prompt tokens (uncached input).
  • Completion tokens (model output).
  • Cache-write tokens (billed at a higher rate by most providers).
  • Cache-read tokens (billed at a discount).
  • Subagent tokens tracked separately from the main agent.
  • Per-model breakdown when the task router mixes providers.

Providers with built-in pricing

Pricing tables ship for the major providers, updated against their public price lists:

ProviderNotes
AnthropicClaude Opus/Sonnet/Haiku with cache-write and cache-read rates
OpenAIGPT-5.4, GPT-4.1, o3, o4-mini
GoogleGemini 2.5 Pro/Flash, Gemini 3 Flash/Pro
DeepSeekV4 (pro/flash), V3, R1; chat + reasoner aliases
GroqLlama 3.3, Llama 4 Scout, Qwen3, GPT-OSS
MistralMistral Large/Medium/Small, Codestral, Magistral, Ministral, Pixtral, Devstral
FireworksTier-based pricing (Mixtral, Llama 70B+, DeepSeek)
GitHub CopilotPremium-request multiplier-based estimation
OpenRouterLive pricing from the catalog
GitHub ModelsPer-token via multipliers
Ollama, LM Studio, OpenCode free models$0.00

Custom providers default to a conservative estimate. Unknown models fall back to Sonnet-tier pricing as a safety floor.

Why it matters

Two tactics cut cost dramatically:

  1. Mix models. Haiku for spark agents, Sonnet for ember agents, Flash for compaction. A task that would cost $0.25 on Sonnet often runs for $0.05 when the exploration phase routes through Haiku.
  1. Use caching. Cache reads are 10x cheaper on Anthropic, up to 50% off on Groq/Fireworks. Empryo structures the system prompt and the Genome for maximum cache hits - typical cache-hit rates exceed 60%.

UI

The status bar shows the running total in USD. Compact mode shows tokens plus a dollar figure. /context opens the detailed view: per-model usage, cache ratio, subagent spend, and the compaction history.

/usage opens the spend ledger — today, this week, this month, per model — plus a cumulative prompt-cache hit rate so you can see caching working across sessions, not just in the current turn. A healthy rate means the cached prefix is holding; a low one means something keeps re-writing it.

Use /router to assign cheap models to cheap tasks.

Flat-rate plans are measured differently

Everything above prices tokens in dollars. Two routes do not bill per token at all — they spend an allowance, and the only limit that matters is how much of it is left:

Those meters sit alongside the dollar figures rather than replacing them, because one session can do both: a gateway lane billing credits and a relay lane burning a weekly cap.

Documentation