Someone asked why the bill tripled. Have the answer by Friday.

Tokenomy sits in the request path as an OpenAI-compatible proxy. Swap the SDK base URL and every call lands in a real ledger with agent, step, tool-call, product and customer attribution. Bring your own keys — provider spend stays on your accounts.

Last updated . Model pricing is refreshed twice daily.

Meter

Every call attributed, not sampled: agent, step, tool call, product, customer, environment, request hash. Attribution at a granularity a billing export structurally cannot reach.

Enforce

Declarative budgets with hard and soft caps, HTTP 402 at quota, z-score anomaly detection and circuit breakers in the request path. A runaway loop stops at 2am instead of appearing on the invoice.

Route

Quality and latency thresholds you define, enforced per request. Response caching with measured hit ROI. The Model Right-Sizing Lab sweeps your workload across models and returns the cheapest passing plan.

Attribute

Upload agent traces or route live. Per-run, per-step and per-tool-call breakdowns show which reasoning loop, retry policy or context payload is burning the budget.

About savings numbers

The waste scanner gives you a modeled percentage with every assumption on screen and editable. The only number worth quoting upstairs is the one measured on your own traffic after the runtime is in place.