TL;DR: Token prices fell roughly 98% since late 2022. Enterprise AI spend still rose about 320% over the same window, according to industry analyses cited by The Next Web and the Linux Foundation's new Tokenomics Foundation. The reason is consumption, not price. Usage-based and hybrid pricing now covers over 60% of SaaS contracts, up from roughly 27% in 2023. If you run a $500K-$5M revenue SaaS business, you cannot manage AI spend with a per-seat mindset anymore. You need a system. I call mine FOCUS: Find every meter, Own the utilization data, Cap the burn rate, Understand the contract terms, Standardize the review cadence. Run it quarterly and you catch the overrun before it hits your P&L, not after.

I spent nine years in the Navy standing watch in engine rooms on gas turbine ships. You do not walk away from the gauges. Ever. Fuel flow, oil pressure, shaft RPM, exhaust temp, you read them on a schedule and act the moment one drifts outside the band. Nobody hands you a monthly invoice explaining what went wrong three weeks after the engine seized. You catch it at the gauge, in real time, or you do not catch it at all.

Most finance teams run AI and SaaS budgets the opposite way. They wait for the invoice. By the time the number lands, the quarter is spent and the only lever left is arguing with a vendor about a bill you cannot itemize. Run your SaaS stack like an engine room, not a mailbox. This article is the research behind why that shift matters, and the five-step FOCUS strategy I use to make it operational.

The Paradox: Cheaper Tokens, Bigger Bills

Here is the number that should worry every operator with an AI line item: GPT-4-equivalent model performance cost roughly $20 per million tokens in late 2022. Today it costs about $0.40 per million tokens. That is a 98% reduction in unit price. By any normal budgeting logic, your AI costs should have collapsed.

They did not. Enterprise AI budgets grew from an average of $1.2 million per year in 2024 to roughly $7 million in 2026, a 320% increase, according to multiple industry analyses tracked by the Linux Foundation's new Tokenomics Foundation. Per-developer token consumption rose 18.6x in nine months, driven almost entirely by agentic tools that do not just answer a question once. They loop: plan, execute, check their own work, retry. Each loop burns tokens the old chat-and-respond model never touched.

This is not a pricing problem. It is a consumption problem, driven by how you deploy AI, not what the vendor charges per token. Uber reportedly burned through its entire 2026 AI budget by April. One company ran up a $500 million Claude bill in a single month. Microsoft revoked developer Claude Code licenses mid-year to bring usage back under control. Seventy-three percent of enterprises report actual AI costs exceeded even their inflated projections.

You do not run at Uber's scale. Neither do I, and neither do most operators reading this. But the mechanism is identical whether your AI line item is $4,000 a month or $4 million. Cheaper unit price plus unmetered usage equals a bigger bill, every time. The gauge that matters is not the price per token. It is the burn rate.

Why Per-Seat Pricing Is Disappearing Under You

For a decade, SaaS budgeting was simple: count heads, multiply by seat price, done. That model is breaking, specifically because of AI.

Consumption-based pricing overtook per-seat pricing as the single most common SaaS pricing model in Q2 2026, at 36.5% of contracts versus 32.4% for per-user pricing, per Vertice's benchmark data. Kyle Poyar's State of B2B Monetization survey puts usage-based adoption at 38% in 2026, up from 27% in 2023, with pure per-seat deals down to just 8% of the market. Sixty-one percent of vendors now run a hybrid model: a base subscription plus a metered layer on top.

The forcing function is what the industry calls the seat apocalypse. AI agents let a 20-person team do the work of 50, while actual workload, tokens processed, records touched, tickets resolved, climbs 10x or more. A per-seat contract charges you less as your team gets more value from a tool. That is backwards, and vendors know it. GitHub Copilot moved 4.7 million subscribers from flat-rate to token billing in June 2026. GitLab built an entirely new consumption model, GitLab Flex, to handle the same uncertainty.

For you as a buyer: 78% of IT leaders were hit with unexpected AI or consumption charges in the past twelve months, and 61% cut planned projects because of price increases tied to that shift, per Zylo's 2026 SaaS pricing trends report. The contract you signed two years ago is not the contract you are paying today. The meter changed underneath it.

The FOCUS Strategy: Five Steps to Control the Burn

I built FOCUS after watching three portfolio-stage clients get blindsided by usage overages in a single quarter. None had a bad vendor relationship. All had a visibility gap. Here is the sequence.

F, Find Every Meter

Pull twelve months of transaction history from every source: corporate card, AP ledger, expense reimbursements, and SSO logs. Do this in parallel, because no single source is complete. Card data misses free-tier tools that later convert. SSO misses anything bought outside identity. Expense reports catch the shadow stack living on personal cards.

The typical mid-market company runs about 112 SaaS applications, and IT knows about roughly 60% of them, per BetterCloud's State of SaaSOps data. Tag every tool by billing type: seat, usage, or hybrid. The usage-billed and hybrid tools are where risk concentrates, because those bills can move without anyone touching a contract.

O, Own the Utilization Data

For every tool over $500 a month, pull actual usage against what you pay for. For seat-based tools, that is active seats divided by contracted seats over a rolling 30-day window. Anything under 60-70% utilization is a downgrade candidate at the next renewal.

For consumption-priced tools, your LLM API bills, your AI coding assistants, your usage-metered analytics platform, the equivalent metric is tokens or credits consumed against your monthly allotment, tracked daily. Vertice's research found spend on consumption-priced tools can vary by as much as 37.6% month to month. If the first time you see that swing is the invoice, you already lost the quarter. You need the daily number, the same way I needed the oil pressure gauge every four hours on watch, not once a month when the maintenance log came due.

C, Cap the Burn Rate

This is the step most operators skip, and the one that actually prevents the overrun instead of documenting it after the fact. Set hard budget thresholds per team, per tool, or per feature, not soft alerts you can ignore, actual caps that stop a request before it lands on the invoice.

Three tiers work for a business your size: a warning at 70% of monthly budget, a hard alert at 90%, and an automatic throttle or block at 100% for anything that is not customer-facing production traffic. If you run AI features through an API, most gateway tools enforce this at the request level, before the model call executes. That is the difference between catching a casualty and reading about one in a report three weeks later.

U, Understand the Contract Terms

Pull the order form and master services agreement for every tool over $1,000 a month. Look for three clauses: renewal notice period, price step-up language, and minimum usage or seat commitments. Most enterprise SaaS auto-renews with a 30-to-90-day notice window. Miss it on a consumption-priced tool with a step-up clause, and you are locked into a higher rate for another year with zero leverage.

Flag anything renewing within 90 days today. The best negotiation conversations happen 60-90 days out, while the vendor still believes you might leave. A simple message to your account rep, saying you are reviewing usage and considering alternatives before renewal, routinely recovers 15-30% off the renewal price. Most account managers hold discount authority they will not use until you ask.

S, Standardize the Review Cadence

A one-time audit finds waste that already accumulated. It does not stop new waste from accumulating next quarter. Build a recurring review: monthly for anything usage-metered, quarterly for the full stack. The monthly review should take under an hour once the pipeline exists: new vendors that appeared, burn rate against budget, any tool trending toward its cap.

Run a full audit before any fundraise, and after any headcount change of 20 people or more, because that is the threshold where seat math and usage tiers shift simultaneously. Most audits I have run surface 15-30% of annual SaaS spend as reclaimable.

What This Looks Like in Practice

A client running a $2.8 million ARR SaaS business came to me convinced their AI coding tool vendor had changed pricing without telling them. What actually happened: the per-token rate had not moved in six months. Their engineering team had adopted an agentic workflow running multi-step autonomous loops instead of single query-response calls. Token consumption per developer rose roughly 4x. The rate card was identical. The bill tripled because the meter spun faster, not because the price per rotation changed.

We ran FOCUS in three weeks. Finding every meter surfaced two overlapping AI tools nobody had reconciled: one team had a departmental subscription, another had individual licenses, both hitting the same model provider. Owning the utilization data showed one was running at 12% of its allotted usage. Capping the burn rate meant setting a hard threshold on the agentic workflows specifically, since that was the actual driver, not blocking the whole tool. Understanding the contract terms found a 60-day renewal notice window closing in 11 days on the larger contract; we caught it with time to negotiate instead of auto-renewing at the old commitment. Standardizing the cadence turned a one-time fire drill into a monthly 45-minute review.

Net result: $34,000 in annual savings, plus a burn-rate cap that has prevented three subsequent months from exceeding budget. None of it required cutting the tool that was actually working. It required seeing the gauge before the engine ran hot.

The Real Lesson: Verification Beats Optimism

Every finance team I have worked with wants to believe the vendor's sales deck: predictable pricing, generous included usage, no surprises. That optimism is exactly how a $1.2 million AI budget becomes a $7 million AI budget without anyone approving the jump in a single meeting.

Verification beats optimism. Not because vendors are dishonest, most are not, but because usage-based pricing shifts the risk from the vendor's cost structure to yours. The vendor's margin is protected either way. Your budget is not, unless you build the instrumentation to see consumption before the invoice does. That is the whole argument for FOCUS: it replaces "I assume this contract is fine" with "I checked the meter this week and here is the number."

FAQ

Why did AI spending grow if token prices fell 98%? Because consumption grew faster than price fell. Per-developer token usage rose roughly 18.6x in nine months as agentic tools replaced single-query interactions with multi-step autonomous loops. Cheaper unit price plus dramatically higher volume produces a bigger total bill, even though each token costs far less than it did in 2022.

Should my SaaS business move to usage-based pricing or stick with per-seat? It depends on whether your marginal cost is flat or variable. Keep seat-based pricing for tools where the end user is a human and cost per user is stable, collaboration and productivity tools. Move to hybrid or usage-based pricing anywhere your cost scales with consumption: any LLM call, API request, or compute cycle behind a feature. About 61% of SaaS vendors now run a hybrid model, balancing predictability with fairness on variable cost.

How often should I audit my SaaS and AI spend? Monthly for anything usage-metered, quarterly for your full stack at minimum. Run an additional full audit before any fundraising round and after any headcount change of 20 people or more, since those events are most likely to break your existing utilization assumptions overnight.

What is the fastest way to find hidden AI or SaaS waste? Pull twelve months of transaction data from your corporate card, AP ledger, expense reports, and SSO logs in parallel, then reconcile into one list. The overlap and gaps between those four sources are where shadow spend and duplicate tools hide. Most first-time audits surface 15-30% of annual spend as reclaimable.

Can I actually cap AI spend before it happens, or only after the invoice arrives? You can cap it before. Modern AI gateways and FinOps tools enforce hard and soft budget thresholds at the request level, meaning a call gets blocked or throttled before the model runs if it would breach your cap. That is meaningfully different from a billing alert that tells you about an overrun after it already happened.

Doctrine Connection

Verification beats optimism. I learned that reading gauges in an engine room, where the cost of trusting your assumptions instead of checking the instrument is a casualty report, not a budget memo. Your AI and SaaS spend deserves the same discipline. Check the meter. Cap the burn. Do not wait for the invoice to tell you what already happened three weeks ago.


*Jeff Barnes, MBA has no personal position in any company, fund, or platform named in this article. demg.ai provides marketing education and consulting services, not investment advice. Past performance does not guarantee future results.*