The Most Useful Line on Your AI Cost Report Is the One You Can’t Explain

Attribution, allocation, and why “unknown” belongs in the schema.

This piece grew out of a comment thread on Sarvar Nadaf’s Per-Agent Cost Tracking for Multi-Agent AI on AWS. The schema below was worked out in that conversation, in public, and it is better for it. Where a specific idea came from the exchange, I have tried to say so.


Most AI cost dashboards answer one question well: how much did this run cost. Total tokens, model spend, per-agent spend, latency, tool usage. Those are real and useful numbers, and for a while they are enough.

They stop being enough the moment your system becomes a composition. Once a request flows through a retriever, a knowledge graph, three specialist agents, and a supervisor that synthesizes their output, the total tells you almost nothing about what to change. A run can be correct, return HTTP 200, look healthy in every latency-and-errors panel, and still cost forty percent more than an identical run that produced the same answer. The overspend is real. It is just not anywhere you are looking.

To find it, you have to stop asking where the money was spent and start asking what caused it to be spent. Those are different questions, and the gap between them is the whole subject of this piece.

Where a Cost Is Incurred Is Not What Caused It

Consider a retrieval operation that costs $0.004: searching, ranking, fetching. That is the direct cost, and it is easy to attribute. It happened on that span, you can measure it, done.

Now suppose that retrieval returned 20,000 tokens, and all of them were hydrated into a supervisor’s context on the next step. The supervisor then costs $0.009. How much did the retrieval really cost?

The direct answer is still $0.004. But that is no longer the interesting answer, because the retrieval also caused cost somewhere else. It inflated the supervisor’s context, and some portion of that $0.009 exists only because the retrieval handed it too much material. The cost was incurred at the supervisor. It was caused, in part, at the retriever.

This gives you two distinct dimensions, and a useful cost model has to carry both:

  • Where the cost was incurred. This is just the span. Directly observed, low ambiguity.
  • What caused or contributed to it. This is the interesting axis, and it is the one no aggregate dashboard shows.

A record that captures both might look like this:

span_id: retrieval-104
direct_cost: $0.0040
primary: retrieval
returned_tokens: 20000

downstream:
  span_id: supervisor-105
  attributed_cost: $0.0021
  caused_by: retrieval-104
  attribution_method: proportional

The dollars are single-counted. We do not charge the $0.0021 twice. The supervisor genuinely incurred it; the retrieval genuinely contributed to causing it; and the record says both without inventing money. What we have added is lineage: a link from a downstream cost back to the decision that helped produce it.

The Hard Part Is Honesty About How You Know

Here is the question that breaks naive versions of this: how do we know retrieval-104 actually caused $0.0021 of the supervisor’s cost, and not some other amount?

Sometimes you can measure it. If you have a controlled comparison where the only meaningful change is that retrieval result, the delta is real evidence. Supervisor costs $0.006 without the retrieved material and $0.009 with it, so roughly $0.003 of downstream cost is attributable to that retrieval. That is measured causation, and it is the strongest claim you can make.

Most production traces do not give you that. In a real run the supervisor is carrying system instructions, conversation state, the outputs of other agents, tool results, and the retrieved material, all at once. There is no clean counterfactual. So you fall back on allocation: split the supervisor’s context cost proportionally by the tokens each source contributed. That is a reasonable method. It is not measurement, and the receipt must not pretend it is.

This is why the single most important field in the whole schema is not a dollar amount. It is this:

attribution_method:
  - measured_delta
  - proportional
  - estimated
  - unknown

That field is what keeps the entire model honest. It stops a proportional guess from masquerading as measured causation. With it, a line can say “retrieval span 104 contributed an estimated $0.0021 of downstream context cost, allocated proportionally by hydrated token share,” and every word in that sentence is defensible, because the method is stated. Without it, the same $0.0021 acquires a precision the evidence never earned.

Resist collapsing this into a confidence score. A number like confidence: 0.82 feels rigorous and gives you nothing, because now you have a second number whose provenance you have to go investigate. measured_delta, proportional, estimated, and unknown each tell you why you are entitled to believe the figure. The method is the provenance. A score would hide it.

Show the Method Where the Decision Is Made

A natural instinct is to keep the attribution method as drill-down metadata, out of the main view, so the report stays clean. That instinct is wrong, and it is wrong for the same reason aggregate dashboards are wrong: it makes two different claims look equivalent.

The method belongs inline, next to any attributed cost, with one sensible exception. A directly observed cost carries no ambiguity and needs no method tag:

retrieval-104   RETRIEVAL   $0.0040

There is nothing to disclose there; it was measured on the span. But the moment a number is attributed rather than observed, the method has to ride along:

retrieval-104 -> downstream CONTEXT   $0.0021   proportional

Drop the word proportional and that $0.0021 visually becomes as solid as the $0.0040 above it, which is a lie of formatting. A report that hides the distinction between what it measured and what it allocated has committed the same sin as the dashboard that only shows a total. If two numbers make materially different claims, the interface must not make them look the same.

So the main report shows amount, category, direct versus downstream, and method. The drill-down holds the evidence behind the method: hydrated token counts, comparison runs, parent-child span references, the assumptions the allocation rests on. The decision surface stays readable; the receipt is one click away, not dumped into the table.

“Unknown” Is Not a Gap in the Accounting

Every honest version of this schema has to allow unknown as a real value, not a placeholder you feel bad about. And once you sit with it, the unknown rows turn out to be the most useful rows in the report.

A high downstream cost with unknown attribution is not incomplete bookkeeping. It is the system telling you exactly where your observability boundary stops letting you explain its own behavior. It is pointing at the place where you cannot yet answer “what caused this,” which is the place most worth instrumenting next. A tidy report with no unknown rows has usually not achieved understanding. It has hidden its ignorance behind confident allocation.

There is also a real reason unknown is sometimes the only honest answer: context is not additive. An extra 5,000 tokens of context does not simply add a proportional slice of cost. It can change caching behavior, alter the reasoning path the model takes, or change how much output the model generates downstream. When that happens, token share and cost share stop mapping to each other cleanly, and any proportional number you report is a polite fiction. In those cases the schema should say unknown and mean it, rather than allocate a figure it cannot defend.

Read that way, the report stops being a statement of where you paid and becomes a map of two things at once: what you can explain about your spending, and where your ability to explain it runs out. The second map is the one that tells you what to build.

What This Actually Costs to Build

The reason this is not a research project is that most of the structure already exists. If your traces are parent-child spans, the causal lineage is physically present already; a downstream cost sits under the decision that produced it. You are not inventing a new tracing mechanism. You are making the attribution semantics explicit on top of a trace you already record.

Concretely, that is a small number of additions. Stamp a primary category on each span. Add a contributes_to link and a cause on spans that produce downstream effects. Add the attribution_method on any attributed cost. Then roll the report up along two axes, category and direct-versus-downstream, and let unknown be a first-class row rather than a swept-under one. The intelligence is not in the plumbing. It is in having the honest method field and being willing to publish the unknown rows.

One warning from the same conversation that produced all this: how you record and how you attribute are coupled. Change the way spans are emitted and you can silently break the logic that reads them. The defense is the same one that makes the whole model trustworthy, a known-good baseline you compare against, so that when your instrumentation shifts under you, the numbers move and you notice.

The Point

We have spent a lot of effort making AI spend visible. Better token counts, per-agent breakdowns, nested traces. All of it answers “how much.” Almost none of it answers “why,” and “why” is the only version of the question you can act on.

The move from one to the other is not a bigger dashboard. It is a small, honest schema: separate where a cost was incurred from what caused it, state the method behind every attributed number, and treat the costs you cannot explain as signal rather than embarrassment. Do that, and the report stops telling you what you spent and starts telling you what to fix, including, in the unknown rows, where to look first.

Even the cost report, it turns out, needs provenance. Not just the amount and the category, but how sure you are about who to blame. That last column may be the most useful one on the page.


With thanks to Sarvar Nadaf, whose post and the conversation under it produced this schema, and to the commenters in that thread who pushed on the baseline and the propagation. The receipt is better for the argument.

Facebooktwitterredditlinkedinmail

Leave a Reply

Your email address will not be published. Required fields are marked *

This site uses Akismet to reduce spam. Learn how your comment data is processed.