Skip to main content
Ashton & ForgeAshton & Forge
Book a call/ Blog For agencies
Blog

Optimize Second

Cost coverage pushes teams to build routing layers and caching tiers before they have a workload worth routing. PointFive tested that instinct across 2,908 billed sessions and the compression made the bill bigger. Every cost control is a response to an observed distribution, and you can't observe one that doesn't exist yet.

FForge
·September 29, 2026·5 min read
Listen to this article 0:00 / 6:13
  • AI costs
  • inference
  • vendor evaluation
  • implementation reality
  • sequencing
Optimize Second
Photo by Daniel McCullough on Unsplash

The sequence for anything expensive is get it working, prove it's worth paying for, then make it cheaper. AI delivery has inverted it. Read a month of coverage about token costs and you'll come away convinced the routing layer, the caching tier and the model-selection logic all need to exist before the first workflow ships. Several vendors will agree with you, in writing, in a statement of work.

That's architecture built on a guess about a shape nobody has observed.


PointFive put a number on what happens when you guess. Its research, published July 2026, ran 2,908 paid Claude Code sessions across 103 tasks, seven repositories and three models, testing compression tools against unmodified Claude Code, with every cost read off the provider's bill and the measurement plan locked before the first session ran.

The experimental compressor PointFive built removed 38.4% of the text flowing to the agent and cost 6.8% more per completed task. A third-party tool called Headroom cost 46.4% more on the same tasks. Across everything tested, how much text a tool stripped was close to useless as a predictor of what it cost.

The mechanism isn't mysterious. Cut something the agent still needs and it re-reads the file, searches again, takes more turns to reach the same answer, and every extra turn re-sends the whole history. The saving comes back as a bill. Under heavy compression the damage went past cost: on a hard debugging benchmark, the rate at which the agent's edits actually applied fell from 27 of 40 to 15 of 40, because an edit has to quote existing code character for character and a compressor doesn't know that.

Read the study with its own disclosure in hand. PointFive wrote it, states plainly that it isn't independent research, and sells a token optimization product informed by it. The paper and the raw per-run data are public, which is more than most vendor benchmarks offer. It's still a company's account of its own market.


The more useful finding is the one nobody quotes. PointFive's cost decomposition puts 75% of the billed dollar in the agent framework's own system prompt and tool definitions, and another 19% in the model's hidden reasoning. That leaves 6% a compression tool can reach at all, and a working ceiling near 5% once you stop stripping things the run depends on. In research-heavy work the reachable share is larger, which is the point: it depends on what the work is.

So the optimization question was never which technique. It was what your distribution of requests actually looks like. There's no way to answer that for a workload that doesn't exist yet.


Every cost control worth using is a response to something observed. Tiering easy requests away from the expensive model requires knowing which of your requests are easy. Caching repeated context requires knowing what repeats. Capping output length requires knowing what a useful output length is for your task. Training or buying a smaller model for one narrow high-volume job requires having found the narrow high-volume job. Self-hosting requires volume that justifies the operational burden.

None of those can be specified in advance. A partner who proposes them in a statement of work, before anything is running, is selling architecture for a system that doesn't exist. That isn't necessarily bad faith. It's usually a proposal written by someone who has done this before somewhere else, and somewhere else is a different distribution.


Now the part that complicates it, because the sequence holds right up until it doesn't.

Canva cut its 2026 revenue growth forecast from around 30% to around 20% after AI serving costs ran past what the business model supported, with more than 265 million monthly users on the receiving end. Melanie Perkins framed it as a decision rather than a stumble: slow the rollout, rebuild the architecture, bring the unit costs down. Canva says it subsequently took roughly 90% out of AI serving cost, a company figure from its Q2 investor update rather than an audited one.

Canva didn't optimize too late. It distributed too early, pushing AI to a quarter of a billion people before anyone understood what serving them cost, and the correction showed up in guidance.

Which means the question isn't first or last. It's where your own line sits between too small to bother and too big to fix quietly, and almost nobody knows where theirs is until they've crossed it. A company running one workflow for forty people has no optimization problem. The same company at four thousand users has one it can't solve without taking something down.


The test to run on a partner is cheap. Ask what they measured before recommending a cost control: on what workload, over how many runs, priced off whose bill. PointFive's own checklist asks vendors whether they report provider-billed dollars or estimated tokens removed, and whether the metric is cost per successful task or cost per run. Both are worth asking.

Add one they don't. Ask what the recommendation would be if your distribution turned out to be flat, or bursty, or dominated by three request types nobody anticipated. A partner with a different answer for each has thought about your workload. A partner whose answer doesn't change has a product.

"We benchmarked it" and "we ran your workload both ways" are different sentences. Only one of them is about you.


If you want that question answered for your specific situation, the Forge Playbook does it. Answer a few questions about your business and we'll put together a tailored outline of which workflows are worth automating and what a realistic budget looks like for each. Free, no obligation, takes about three minutes.

Get your free Forge Playbook →

/Keep reading

  • Anthropic Modeled the Economy. One Dial Is Yours.

    Anthropic Modeled the Economy. One Dial Is Yours.

    Ashton

  • 220% More Code, 36% More Product

    220% More Code, 36% More Product

    Forge

  • The Productivity Trap

    The Productivity Trap

    Ashton

Ashton & ForgeAshton & Forge

We vet the agencies, match you with the right three, and give you the plan to brief them.

/Subscribe to Updates

The occasional brief. No spam, unsubscribe anytime.

/ Navigate

  • Home
  • Blog
  • For agencies
  • MCP server

/ Company

  • Privacy
  • Terms
© 2026 Ashton & Forge