Home
A Kill‑A‑Watt meter for your AI agents.
Wattage reads OpenTelemetry GenAI semantic-convention traces — the standard most agent frameworks and observability tools already emit — and tells you exactly where your tokens are being burned and wasted. It quantifies the dollar cost of each waste pattern, prescribes a fix, computes a quality-aware Token Efficiency score, and can fail your CI when a change makes your agents more expensive.
What it is (and isn't)
Wattage is a diagnosis + prescription + gate, not another dashboard:
- It consumes the traces your existing observability tool (Langfuse, Helicone, an OTel Collector, whatever) already produces — it doesn't replace them.
- It names specific waste patterns with dollar figures and a fix, rather than just showing you a bill.
- It never enforces anything automatically — findings are recommendations you apply, not prompts it silently rewrites. Runtime enforcement (killing waste in-flight) is a deliberately separate, later, opt-in capability.
The three surfaces
wattage report— a priced, findings-quantified report for one trace: terminal, JSON, or a self-contained HTML flame graph.wattage score/wattage badge— a single 0–100 Token Efficiency grade you can drop into a README.wattage ci— a cost-regression gate: fails a PR when a change makes your agent measurably more expensive, with exact, documented exit codes.
Quick start
uvx wattage report trace.json
That's it — no config file, no API key, fully offline. Point it at an OTLP JSON export of your agent's trace and it prices every call, runs every detector, and prints a report. Don't have one yet? See Getting your first trace — covers both "I already have OTel traces" and "I have zero instrumentation."
Where to go next
- Getting your first trace — start here if you don't already have an OTLP JSON trace to point Wattage at.
- Adapters — what trace formats Wattage reads, and how it tolerates the real-world naming differences between frameworks.
- Detectors — the eight waste patterns Wattage looks for, one page each.
- The Convergence Engine — the standout: how Wattage catches agents thrashing in unproductive loops, and why exact-match duplicate detection can't.
- CI Integration — wiring
wattage ciinto a GitHub Action so cost regressions fail the build.
A note on honesty
Every dollar figure, every score, and every "typical savings" claim Wattage produces comes from an actual computation against your real trace and the current vendored pricing data — never a guess dressed up as a number. When Wattage doesn't have enough information to say something confidently (an unpriced model, an unmeasured quality signal, a genuinely ambiguous loop), it says so explicitly instead of filling the gap with a plausible-looking number. That's a design principle, not an afterthought — see the individual detector pages for exactly which limitations each one is honest about.