Token Economics: When Your SaaS Bill Stops Counting Seats

Nature uses tokens as well ;-)

A lot of SaaS pricing still assumes the same thing: value scales with the number of seats. More people in the tool? More revenue. Simple.

But this shift isn’t limited to “SaaS pricing”. You see the same consumption logic across cloud and data services that don’t bill per VM-hour, but still behave like tokens (think Microsoft Fabric).

And a token-based model doesn’t care how many people have access. It cares how much the system is used. That’s great for adoption — and often bad for your wallet — because usage tends to scale faster than headcount.

This piece is about that change: tokens vs. seats, what it breaks, what it enables, and what you should measure before your next "great deal" becomes a recurring mistake.

The old mental model: seats as a proxy for value

Seat-based pricing works because it is easy:

  • predictability for finance

  • procurement-friendly unit pricing

  • clear limits for access control

In practice, seats are rarely a clean proxy for business-derived value. They are a proxy for potential usage.

If ten people have access, but two actually use the tool, you're paying for organisational intent, not outcome. That mismatch exists in every enterprise portfolio, but it stayed manageable because the unit (a seat) is slow-moving: headcount changes monthly, not hourly.

Tokens change the unit: from identity to consumption

Tokens flip the unit from who to how much.

That has two immediate consequences.

1) Usage becomes the cost driver

The bill follows consumption. Not adoption. Not enablement. Actual interaction volume.

If you have never measured usage beyond "active users", you now have a problem: activity is not the same as consumption. In token models, one highly active workflow can be more expensive than a whole department occasionally clicking around.

The bigger problem: most organisations don’t have the instrumentation to judge whether that consumption is “good” or “waste”. If you can’t connect tokens to an outcome, you’re blind — and you’re paying to stay blind.

2) Cost becomes elastic — in both directions

Elasticity is the promise. You pay less when you use less.

But elasticity is also the risk: the cost can scale without warning when:

  • a workflow gets automated

  • a feature is rolled out broadly

  • a new system starts calling an API at scale

  • an agent-loop runs overnight because nobody set guardrails

  • an agent that orchestrates other agents with poor context and sloppy prompts, creating expensive rework at scale

Tokens don't create that risk. They expose it.

Token economics in the "post–gold rush" reality

There was a brief gold rush phase where tokens felt like a discount. Early pricing, hype cycles, "experimental" budgets — plenty of organisations treated token spend as noise. Remember Tokenmaxxing? Joke of the year!

That phase is ending.

Tokens are becoming a standard billing unit because they map to a real underlying cost driver: compute and throughput. And that means the model will get tighter, not looser.

If you want to steer this properly, stop treating tokens as "AI spend" and start treating them as a consumption metric — like GB, vCPU-hours, or requests.

The gold rush narrative is real — and mostly financial. Vendors are pouring billions into AI and buying time with investor expectations. That time runs out. When shareholders start demanding green numbers (and a believable growth curve), pricing models tighten. Tokens are one way to make “compute and throughput” show up directly in the bill — and to protect margin.

Why seat vs token is a governance question, not a price question

Most procurement discussions around tokens focus on: "Is it cheaper than seats?"

That's the wrong first question.

The first question is: "Do we have governance for variable consumption?"

Because a token model is, effectively, a mini-cloud inside your SaaS contract:

  • usage spikes

  • unit economics

  • hidden coupling to engineering changes

  • budget predictability challenges

In other words: this is FinOps territory, even when it sits on a software line item. you cannot control what you cannot see — and tokens make that visible in real time.

FinOps is not a cost-saving initiative. It is a governance discipline.

What you should measure (before you renegotiate)

If your organisation is moving from seats to tokens (or mixing both), you need a baseline. Not a spreadsheet guess.

At minimum, measure these four things.

1) The consumption curve

Not the monthly total. The shape.

Where do spikes happen? Are they human-driven (office hours) or system-driven (nightly jobs)? If you can't answer that, you can't set guardrails.

2) Cost per outcome

Some teams now frame this as a KPI: cost per valued outcome (CPVO). That’s the right instinct: not “how many tokens did we burn?”, but “what value did we generate out of these tokens?”

Token cost is useless without an outcome metric.

Pick one outcome per workflow. Examples:

  • cost per generated customer response

  • cost per document summarised

  • cost per ticket triaged

  • cost per policy check executed

If you can't define an outcome, you don't have a business case, you have an experiment.

3) Who (or what) consumes

Token spend often isn't a person. It's:

  • an integration

  • an automated workflow

  • an agent

  • a system account

If you only track "users", you will miss 80% of your spend.

4) Guardrails and failure modes

Set explicit limits:

  • rate limits

  • budget caps

  • alerting thresholds

  • fallback behaviour when tokens run out

This is the difference between "variable cost" and "unbounded cost".

A practical comparison: seats are stable, tokens are steerable

Seat models optimise for stability:

  • predictable invoices

  • easy allocation

  • slow drift

Token models optimise for steerability:

  • pay for actual usage

  • direct coupling between design choices and cost

  • visibility into what drives spend — if you instrument it

The trade-off is clear: you're swapping procurement simplicity for operational responsibility.

And that is not inherently bad. It's just a different operating model.

Also: expect hybrid models. GitHub Copilot is a good example — you pay a seat, plus variable usage on top.

The takeaway

If you treat token spend as a new subscription type, you'll be surprised — and usually not in a good way.

If you treat it as consumption, you can manage it like everything else you already manage in cloud economics: visibility first, then governance, then optimisation.

One question to end on: if your token bill doubled next month, would you be able to explain why within one day?

FAQ

Next
Next

You Cannot Control What You Cannot See — The Origin Story