Per-seat AI budgeting can’t see the tail

Per-seat AI budgeting can’t see the tail

Revenium CTO and co-founder John D’Emic says hat per-seat budgeting can obscure the extreme consumption patterns driving AI costs, leaving organisations exposed to expensive outlier sessions that conventional forecasting and monthly spending controls may fail to detect.

In May, one of our developers opened an AI coding session on his laptop and left it running. The session stayed open for days, making 4,819 calls at a cost of US$3,762 by the time it closed.

Our budget did not anticipate this, nor did an alert fire because nothing ‘broke’. Every one of those calls was metered, so measurement wasn’t the missing piece. But what was missing was a rule that acted on what the data already showed.

We then went back and pulled out three months’ worth of our engineering telemetry to try and figure out how unusual that session was. It wasn’t. More concerning, the shape of the data suggested the way most orgs budget AI right now isn’t able to see the part of the curve where the money truly goes.

Averages obscure

Across 557 code-implementation tasks over those 90 days, our median cost was US$2.24. Our most expensive single task was about US$300, or roughly 134x the median and 65x the average. The four-day session clocked in at US$3,762.

When we widened the aperture, the distribution got starker. Of our 14,680 AI runs, the top 1% accounted for nearly half of our total spend. The top 5% was close to 80% of our spend. That’s a power law, which is not the distribution that typical cost controls are built for.

The FinOps Foundation’s guidance on token economics gets this right and says so directly by naming runaway agentic loops and context window accumulation in long sessions (among other anomaly patterns to watch out for). What our data adds is a sense of scale, meaning how far outside the ordinary these events can land once you have 90 days of traces to look at.

The same guidance recommends instrumenting and measuring for 30-60 days to establish a baseline. From there, the thinking is you can set budgets somewhere around 110-120% of that baseline. Good advice for spend that clusters near the mean, but against a distribution like ours (and, in my experience talking to folks at Ai4 and elsewhere recently, many others), a monthly aggregate of 120% of baseline has enough headroom to swallow up one runaway session inside an ordinary team variance and never trip. The threshold, while accurate, points to the body of the curve as the tail writes most of the check.

At the core, the bill is the people

Current conversation about runaway AI cost is understandably focused on agents, and specifically those autonomous loops retrying until something gives. That produces headlines like teams waking up to a US$47,000 bill (or much more) after four agents spent eleven days talking to one another. It’s a well-founded worry, with Carnegie Mellon finding that even capable agents finish real-world office tasks on their own slightly less than one-third of the time.

From our own data, over three months the automated pipeline that implements and reviews our pull requests generated 4,171 traces and US$6,723, or about 6% of our AI spend. Interactive sessions (our engineers working through the day in Claude Code, Cursor, Codex and tools like them) generated 10,005 traces and US$109,118, or about 94% of the bill.

Tokenomics Foundation guidance sorts AI purchasing into five procurement models, with AI developer tools as one of the easier categories to govern, because seat-based billing means the vendor mediates the API call and the invoice behaves like any other per-seat contract. However, that can hold on the attribution invoice while falling apart on consumption.

What our numbers showed was that costs can be plenty attributable while very hard to forecast (and almost the entirety of our bill). That category also doesn’t tend to get its own line item, as AI coding assistants often sit inside an engineering software budget next to IDEs and CI minutes. Those are costs that stay roughly where you put them from one quarter to the next.

Seat count stopped predicting anything

In January, seven of our engineers consumed US$109 of API-equivalent value between them; by February that count was 11 engineers and US$4,760; by March it was 21 engineers and US$14,463; by April, 27 engineers and US$26,369. By May, engineers consumed US$45,728.

Headcount roughly quadrupled while consumption went up 420x, so call it a hundredfold increase per engineer. Point being that a forecast built in January based around seat count would have been wrong by two orders of magnitude by May. More concerning, it would have looked correct the entire time, because the seat count tracked the prediction pretty closely.

The variable moving was the depth of use per person, and a per-seat model has nowhere to put that variable. Adoption curves and consumption curves came apart in February, never to reconverge.

This isn’t specific to us, and it’s a growing concern that is starting to bite more orgs. Those rolling out AI dev tools are watching this happen as engineers get more fluent with the tools and managers ask bigger things of them. Fluency is, of course, the point, but it can be a massive cost driver that’s largely invisible to a budget structured primarily around seats.

Reading this honestly

These figures come from one engineering organization of about 30 people, instrumented by the product we sell. I’m reporting our own numbers rather than aggregated customer data because they are auditable and attributable to a specific company. But one engineering team is a sample of one.

I would claim the shape of the distribution generalizes, though, since it tracks with what we are seeing and hearing from others in similar situations. But I would not claim our specific numbers are anyone else’s.

But the part I keep coming back to is where this turned up. We build AI cost tools and of course instrument the same product we sell (and this problem is what our teams think about every working day). We still ran a four-day session that no one noticed until we went looking. An organization with none of that has the same distribution running underneath it right now and no view at all.

What would have caught it

Break interactive AI out as its own line item rather than folding it into engineering software spend. If it turns out to be 90%+ of your AI bill, you’ll certainly want to have been watching that from the start.

Treat seat count as a licensing fact but not as a forecasting input. Forecast against consumption per engineer while tracking how that number moves month-over-month.

Set ceilings at the run and session levels, alongside the monthly aggregate. Expensive outliers are individual executions, and a monthly number isn’t granular enough to catch a single execution.

Alerting on session duration and call count next to dollars would have flagged our May 13 session days earlier because 4,819 across four days is a clear anomaly in call volume long before it’s a clear spend anomaly.

To be clear, nothing our engineer did was wrong during that May session. He was doing his job and left a window open the way we all do with tools we rely on. The assumption sitting underneath the budget is the variable that mattered, which was that a seat is a sensible unit for something whose cost is set by how deeply one person decides to use it.

Browse our latest issue

Intelligent CIO North America

View Magazine Archive