Your AI agent budget is not a forecasting problem. It is a live casualty waiting to happen. A San Francisco startup called Polsia watched its revenue climb from $100,000 to a $10 million run rate this year. Its Anthropic token bill climbed right alongside it, hitting $1.2 million a month before anyone in the engine room noticed the flooding. That is not a growth story. That is a company taking on water while the crew celebrates the speed.
I spent years standing watch in the engine room of a Navy nuclear submarine. You learn one lesson fast down there: the systems that kill you are never the ones you're watching. They're the ones running quietly in the background, compounding, until someone finally checks the gauge and finds the number has already left the dial. Inference cost is that gauge for every AI-agent company right now. Most founders aren't checking it. They're checking growth.
Direct answer: Your AI agent budget is a ticking time bomb because token costs scale with usage, not with revenue discipline, and most companies have no metering, no routing, and no ownership of the infrastructure running their agents. The fix is not a smaller model. The fix is sovereignty over the inference layer itself, through routing, ownership, or both, before the bill outruns the business.
The Bomb Is Already Ticking
Polsia's story isn't a one-off horror story. It's a preview. The company runs swarms of AI agents to operate other businesses, employing effectively nobody as a human staff. That's the pitch. Agents replace headcount. But agents don't replace cost. They relocate it, from payroll to tokens, and tokens don't take vacation, don't get laid off, and don't stop billing when the model calls a tool it didn't need to call.
Sapiom, the startup Polsia hired to fix the problem, cut that $1.2 million monthly bill roughly tenfold, down to about $100,000, through model routing and inference optimization (Semafor). Same workload. Same agents. One-tenth the cost. That gap between what companies pay and what they need to pay isn't a rounding error. It's the entire margin of an early-stage company, sitting unclaimed on the table because nobody built a system to claim it.
Sapiom's founder and CEO, Ilan Zerbib, put it bluntly: "It's just unsustainable... startups can't deploy their product at the costs that frontier-model providers are charging" (Semafor). He's not describing an edge case. He's describing the default state of every agent-heavy company that hasn't built cost governance into its architecture.
Systems Beat Slogans
Here's the doctrine, and it applies whether you're running a submarine reactor or a customer support agent fleet: systems beat slogans. "Move fast" is a slogan. A routing layer that sends 95% of your calls to a cheaper model and reserves the frontier model for the 5% that actually needs it is a system. Zerbib says exactly that ratio out loud: "In 95% of cases, it doesn't make sense to go to a very expensive frontier model" (Semafor).
Most founders don't know that ratio for their own stack. They don't know it because nobody measured it. That's not a technology gap. That's a governance gap, and governance gaps always show up on the balance sheet before they show up in the postmortem.
Naive Inc, a startup raising capital to automate the grunt work of running a company, hit the same wall from a different angle. Building serverless infrastructure for autonomous businesses, the company discovered that inference has become the single largest cost line for agent-heavy startups, ahead of cloud compute, ahead of tooling, ahead of everything else on the P&L (TechCrunch). Naive raised $28.5 million and is pouring a meaningful share of it into a model router, a memory layer, and a lightweight serverless runtime built specifically to keep agents from bleeding tokens while they sit idle.
Two independently funded companies, working the same problem from opposite directions, arrived at the identical conclusion: unmanaged inference spend is the new operational risk. That's not a coincidence. That's a market discovering, in real time, that nobody built the damage-control manual for this compartment of the ship.
Naive's CEO, Sean Dorje, said it plainly: inference is now the biggest cost line for companies running agents, ahead of cloud compute, ahead of everything else on the ledger (TechCrunch). That single sentence should reorder every board deck in this category. Founders pitch investors on unit economics for headcount replacement. Few pitch a plan for what happens when the replacement labor's cost structure behaves like a variable-rate loan instead of a fixed salary.
Uber gave the corporate world a live-fire demonstration of that variable-rate loan going wrong. The company burned through its entire 2026 AI budget in four months, and COO Andrew Macdonald admitted the surge in token spend on coding tools has no clear link to shipped features: "That link is not there yet" (The Verge). A company that size, with finance teams and forecasting discipline most startups can only envy, still got surprised by its own consumption curve. That should worry every founder who assumes their smaller operation has better visibility.
Compartmentalize or Sink
On a submarine, you don't let one flooded compartment take down the whole vessel. You seal it off. You compartmentalize the damage, run the casualty drill, and keep the boat afloat while you fix the leak. Most AI-agent companies today have zero compartmentalization on their inference spend. One overactive agent, one bad prompt loop, one integration that calls the frontier model when a smaller model would do, and the flooding spreads through the entire budget with no bulkhead to stop it.
This is where The Sovereignty Stack comes in. The doctrine is simple: own your infrastructure, or someone else owns your margin. Sapiom didn't just build a router. It built its own server racks in San Jose, running open-weight models directly, so it can charge customers for compute at cost, with no markup layered on top (Semafor). That's sovereignty. That's the difference between renting your engine room from a landlord who can raise the rate whenever demand spikes, and owning the hull outright.
Compare that to the alternative most companies default to: a single frontier-model contract, no routing, no fallback, no visibility into which calls actually need premium reasoning and which don't. That's not a technology stack. That's a hostage situation with a monthly invoice.
The market has noticed. OpenRouter's weekly token volume went from 5 trillion to 25 trillion in six months flat, on its way to a quadrillion tokens processed this year (OpenRouter). Amazon and Microsoft have bundled intelligent prompt routing directly into Bedrock and Azure. One market tracker counts more than 80 active routing competitors fighting for the same commoditizing layer (Semafor). When a category attracts 80 competitors and 5x volume growth in half a year, it tells you two things. The pain is real, and the solution is becoming table stakes, not a differentiator. If routing is table stakes, the companies without it aren't behind the curve. They're off the map entirely.
The Payback Period Nobody Calculates
Every capital allocator I've worked with, going back to my time underwriting risk at Hartford and later at Munich Re, asks the same question before committing a dollar: what's the payback period, and what's the downside if the assumption breaks? Nobody asks that question about their inference bill. They treat it like a utility, a fixed cost of doing business, when it behaves more like a used position that can move against you overnight.
Do the math the way an underwriter would. If your agent fleet runs on frontier-model pricing exclusively, and your usage scales 100x in a year the way Polsia's did, your inference bill scales close to 100x too, unless you've built a system to intercept it. That's not a growth story. That's an unhedged bet on a single vendor's pricing model, sized at whatever your growth curve decides to make it.
A KPMG Global AI Pulse survey of more than 2,145 senior leaders in June found only 7% could point to established returns from their AI spending (KPMG). Forrester predicted enterprises will defer a quarter of planned AI spending into 2027, as CFOs demand ROI before approving further budget (Forrester). The corporate world is running its first real audit of AI ROI, and the audit is not going well for companies that never built a metering system in the first place.
This is due diligence, plain and simple, and due diligence is non-negotiable. You would never underwrite a policy without knowing your loss ratio. You should never scale an agent fleet without knowing your cost-per-task, your routing ratio, and your ceiling before the invoice arrives.
What the Doctrine Requires
The fix isn't complicated. It's disciplined. Three moves, in order.
None of these moves require exotic technology. They require the same posture I learned standing reactor watch: assume the leak is already happening, and build the instrumentation to find it before the crew smells smoke. Waiting for the invoice to tell you the story means you've already lost a month of runway you can't get back.
First, meter everything. You cannot compartmentalize a leak you cannot see. Instrument every agent call, every model invocation, every token, before you scale headcount-replacement further. Naive built its governance layer specifically to enforce budgets and require human approval before agents burn spend unsupervised (TechCrunch). That's not bureaucracy. That's a bulkhead.
Second, route by task, not by habit. The 95% rule Zerbib cites isn't theoretical. Most agent tasks are simple classification, retrieval, or formatting work that a smaller, cheaper model handles at a fraction of the cost. Reserve the frontier model for the 5% of calls that actually require frontier reasoning. Everything else is waste dressed up as convenience.
Third, decide whether you rent or own. Renting infrastructure through a router is faster to deploy. Owning infrastructure, the way Sapiom does with its own San Jose server racks, is slower to stand up but removes a landlord from your cost structure permanently. Neither choice is wrong. Not choosing is wrong. That's the asset question every board should be asking: is our inference layer an asset we control, or a liability we've outsourced blind?
The exit for this problem doesn't come from cutting agents or slowing growth. It comes from turning inference from an unmanaged liability into a governed line item with a known payback period. Companies that do this now will scale their agent fleets on a stable foundation. Companies that don't will discover their bomb the way Polsia did, after the fuse is already lit.
FAQ
Why did Polsia's AI costs spike so fast? Its agent-driven revenue grew nearly 100x in a year, from a $100,000 to a $10 million annual run rate, and its Anthropic token usage scaled right alongside it, with no routing or optimization layer in place to slow the climb (Semafor).
What is model routing, in plain terms? It's a system that sends each AI task to the cheapest model capable of handling it correctly, reserving expensive frontier models only for tasks that genuinely require that level of reasoning, rather than sending every call to the priciest option by default.
Is model routing enough, or do I need to own infrastructure too? Routing alone can cut costs dramatically, as Sapiom's tenfold reduction for Polsia shows. Ownership goes further by removing markup entirely, since Sapiom runs open-weight models on its own hardware and charges compute at cost (Semafor).
Is inference cost really the biggest line item for agent-heavy startups? According to Naive's CEO, yes. Inference has become the largest recurring cost for companies running autonomous agents, ahead of traditional cloud infrastructure spend (TechCrunch).
Is routing a durable advantage or a temporary edge? It's becoming a commodity fast. Amazon and Microsoft now bundle routing into Bedrock and Azure, and more than 80 competitors are active in the space. The durable advantage shifts to whoever owns infrastructure and governance, not whoever simply resells routing.
The gauge is right there. Check it before the compartment floods.
*Jeff Barnes is the founder of DEMG.ai and has no personal financial position in any company, fund, or platform named in this article. DEMG.ai provides marketing education and systems for owner-operators, not investment advice. Past performance does not guarantee future results.*