Ashton & ForgeAshton & Forge
Blog

The Token Bill You Didn't Budget For

The licensing number your vendor quoted is real. It's also not the cost. Here's how inference costs compound — and what to ask before you sign.

FForge
··4 min read
Listen to this article 0:00 / 5:24
  • budgeting
  • AI costs
  • vendor evaluation
  • inference
  • mid-market
The Token Bill You Didn't Budget For
Photo by Jon Moore on Unsplash

Most AI vendor conversations start with a licensing number. That number is real, but it's not the cost. The cost is licensing plus inference, and the two don't scale together.

Licensing is the flat fee: the platform subscription, the enterprise agreement, the per-seat charge. Inference is what you pay for the model to actually run, and it's priced per token — every word in, every word out, every call to the API. A $50,000 annual platform license can sit next to a $200,000 annual inference bill if the workflows are high-volume and nobody modeled it out before go-live.

The gap between those two numbers is where most mid-market AI budgets break.

Why the pilot number lies

Inference costs compound in ways that aren't obvious from a static pilot.

In a pilot, you test one workflow on a contained dataset with a handful of users. The model runs a few thousand calls, the results look good, the costs look fine. What the pilot doesn't show you is what happens when you expand. Adding a second workflow doesn't double inference costs — it might triple them, because the new workflow has longer prompts, more context, or higher call frequency. Adding users multiplies throughput. Adding more data into the model's context window multiplies cost per call.

The compounding works like this: longer inputs cost more per call. More calls per workflow add up linearly. Multiple workflows stack. Any retrieval layer that feeds documents into context adds its own token overhead on top of the generation cost. If your use case involves reasoning over long documents, you can easily hit 20,000–50,000 tokens per transaction. At current market rates, that's $0.06–$0.30 per call depending on the model, before you've touched the licensing line.

Run a hundred of those calls a day and it's a real number. Run a thousand and you're replacing a salaried employee's annual cost every few weeks.

How to model the bill before you sign

Responsible cost modeling before you sign isn't complicated, but vendors rarely offer it unprompted.

Start with a unit economics estimate: what's the average input length per call for this workflow, what's the expected output length, how many calls per user per day, and how many users. Use those numbers to calculate a daily token volume. Apply the vendor's published per-token price. Then apply a 2x buffer for prompt overhead, error handling, and the fact that actual usage almost always runs above the estimate in production.

That buffer isn't conservative padding. It's what happens when a workflow that was modeled on clean, structured inputs meets actual user behavior and actual document variation.

Then build a scaling scenario. What does the model look like at 3x the initial user count? At 5x workflows? If the math breaks down before you hit the scale you're planning for, you need to know that before you sign the contract, not six months into a committed annual agreement.

One thing to verify: whether inference costs are included in the platform license or billed separately. Some vendors bundle usage up to a ceiling, then charge overages. Some price them entirely separately. Some enterprise agreements fix a token allocation that sounds generous in a pilot and constrains you at production volume. Read the pricing schedule, not just the headline number.

Four questions to ask any vendor

Before any AI vendor conversation gets to contract stage, four questions cut through the pricing complexity.

First, ask for a worked example of inference costs at scale — not a generic pricing table, a specific scenario at your expected volume. If the vendor can't produce one, that's a problem.

Second, ask where the cost ceiling is in the contract. Fixed-fee agreements with usage caps can be fine if the cap is modeled correctly. Uncapped inference billing against a fixed workflow budget is a risk that will surface eventually.

Third, ask what happens when you add a workflow. Not "what's the price of additional seats" — ask what happens to the total inference bill when you expand scope. This is the question that surfaces whether anyone has actually modeled your deployment.

Fourth, ask for the last three months of billing data from a comparable client. Reference architectures in sales decks are built for the demo; live client billing reflects what the model actually costs to run at production volume.

If those questions produce vague answers, the vendor's incentives are not aligned with your ability to plan. That's not a deal-breaker on its own, but it's the kind of mismatch that produces surprises in month eight.


If you want that question answered for your specific situation, the Forge Playbook does it. Answer a few questions about your business and we'll put together a tailored outline of which workflows are worth automating and what a realistic budget looks like for each. Free, no obligation, takes about three minutes.

Get your free Forge Playbook →

Ashton & ForgeAshton & Forge

We vet the agencies, match you with the right three, and give you the plan to brief them.

/Subscribe to Updates

The occasional brief. No spam, unsubscribe anytime.

© 2026 Ashton & Forge