Token Economics: When Your SaaS Bill Stops Counting Seats
Nature uses tokens as well ;-)
A lot of SaaS pricing still assumes the same thing: value scales with the number of seats. More people in the tool? More revenue. Simple.
But this shift isn’t limited to “SaaS pricing”. You see the same consumption logic across cloud and data services that don’t bill per VM-hour, but still behave like tokens (think Microsoft Fabric).
And a token-based model doesn’t care how many people have access. It cares how much the system is used. That’s great for adoption — and often bad for your wallet — because usage tends to scale faster than headcount.
This piece is about that change: tokens vs. seats, what it breaks, what it enables, and what you should measure before your next "great deal" becomes a recurring mistake.
The old mental model: seats as a proxy for value
Seat-based pricing works because it is easy:
predictability for finance
procurement-friendly unit pricing
clear limits for access control
In practice, seats are rarely a clean proxy for business-derived value. They are a proxy for potential usage.
If ten people have access, but two actually use the tool, you're paying for organisational intent, not outcome. That mismatch exists in every enterprise portfolio, but it stayed manageable because the unit (a seat) is slow-moving: headcount changes monthly, not hourly.
Tokens change the unit: from identity to consumption
Tokens flip the unit from who to how much.
That has two immediate consequences.
1) Usage becomes the cost driver
The bill follows consumption. Not adoption. Not enablement. Actual interaction volume.
If you have never measured usage beyond "active users", you now have a problem: activity is not the same as consumption. In token models, one highly active workflow can be more expensive than a whole department occasionally clicking around.
The bigger problem: most organisations don’t have the instrumentation to judge whether that consumption is “good” or “waste”. If you can’t connect tokens to an outcome, you’re blind — and you’re paying to stay blind.
2) Cost becomes elastic — in both directions
Elasticity is the promise. You pay less when you use less.
But elasticity is also the risk: the cost can scale without warning when:
a workflow gets automated
a feature is rolled out broadly
a new system starts calling an API at scale
an agent-loop runs overnight because nobody set guardrails
an agent that orchestrates other agents with poor context and sloppy prompts, creating expensive rework at scale
Tokens don't create that risk. They expose it.
Token economics in the "post–gold rush" reality
There was a brief gold rush phase where tokens felt like a discount. Early pricing, hype cycles, "experimental" budgets — plenty of organisations treated token spend as noise. Remember Tokenmaxxing? Joke of the year!
That phase is ending.
Tokens are becoming a standard billing unit because they map to a real underlying cost driver: compute and throughput. And that means the model will get tighter, not looser.
If you want to steer this properly, stop treating tokens as "AI spend" and start treating them as a consumption metric — like GB, vCPU-hours, or requests.
The gold rush narrative is real — and mostly financial. Vendors are pouring billions into AI and buying time with investor expectations. That time runs out. When shareholders start demanding green numbers (and a believable growth curve), pricing models tighten. Tokens are one way to make “compute and throughput” show up directly in the bill — and to protect margin.
Why seat vs token is a governance question, not a price question
Most procurement discussions around tokens focus on: "Is it cheaper than seats?"
That's the wrong first question.
The first question is: "Do we have governance for variable consumption?"
Because a token model is, effectively, a mini-cloud inside your SaaS contract:
usage spikes
unit economics
hidden coupling to engineering changes
budget predictability challenges
In other words: this is FinOps territory, even when it sits on a software line item. you cannot control what you cannot see — and tokens make that visible in real time.
FinOps is not a cost-saving initiative. It is a governance discipline.
What you should measure (before you renegotiate)
If your organisation is moving from seats to tokens (or mixing both), you need a baseline. Not a spreadsheet guess.
At minimum, measure these four things.
1) The consumption curve
Not the monthly total. The shape.
Where do spikes happen? Are they human-driven (office hours) or system-driven (nightly jobs)? If you can't answer that, you can't set guardrails.
2) Cost per outcome
Some teams now frame this as a KPI: cost per valued outcome (CPVO). That’s the right instinct: not “how many tokens did we burn?”, but “what value did we generate out of these tokens?”
Token cost is useless without an outcome metric.
Pick one outcome per workflow. Examples:
cost per generated customer response
cost per document summarised
cost per ticket triaged
cost per policy check executed
If you can't define an outcome, you don't have a business case, you have an experiment.
3) Who (or what) consumes
Token spend often isn't a person. It's:
an integration
an automated workflow
an agent
a system account
If you only track "users", you will miss 80% of your spend.
4) Guardrails and failure modes
Set explicit limits:
rate limits
budget caps
alerting thresholds
fallback behaviour when tokens run out
This is the difference between "variable cost" and "unbounded cost".
A practical comparison: seats are stable, tokens are steerable
Seat models optimise for stability:
predictable invoices
easy allocation
slow drift
Token models optimise for steerability:
pay for actual usage
direct coupling between design choices and cost
visibility into what drives spend — if you instrument it
The trade-off is clear: you're swapping procurement simplicity for operational responsibility.
And that is not inherently bad. It's just a different operating model.
Also: expect hybrid models. GitHub Copilot is a good example — you pay a seat, plus variable usage on top.
The takeaway
If you treat token spend as a new subscription type, you'll be surprised — and usually not in a good way.
If you treat it as consumption, you can manage it like everything else you already manage in cloud economics: visibility first, then governance, then optimisation.
One question to end on: if your token bill doubled next month, would you be able to explain why within one day?
FAQ
-
Token economics refers to the financial models that govern how software is priced and billed based on consumption (tokens) rather than user seats. Unlike seat-based pricing, token-based billing charges you for actual usage — the more the system is used, the higher the bill.
-
Measure the consumption curve (spikes and patterns), cost per business outcome, the actual consumers (people, agents, integrations, not just "active users"), and explicit guardrails (rate limits, budget caps, alerting). This is token economics applied as governance.
-
Because tokens create elastic, variable costs that can scale silently when workflows change, features roll out, or automation runs unchecked. You cannot control what you cannot see — so governance, visibility, and tagging become essential, just as in cloud economics.
-
Seats are a proxy for potential usage (stable, predictable, slow-moving); tokens are a metric for actual usage (elastic, real-time, variable). Seats optimise for procurement simplicity; tokens optimise for operational steerability — at the cost of requiring active governance.
