Same Model, Up to Five Times the Bill
Berkeley and Arena held the model constant across three agent harnesses and found similar success rates at up to five times the cost. Uber shows what it looks like when someone treats that spread as an engineering variable. Questions to ask a partner about the part of your bill they actually control.
- AI costs
- vendor evaluation
- implementation reality
- agentic development
- unit economics
On 16 September, researchers at UC Berkeley and Arena published a study that held the model still and changed only the software wrapped around it. Seven models from Anthropic and OpenAI, three agent harnesses (Claude Code, Codex CLI and the open-source Pi), thirty randomly sampled tasks from each of two benchmarks, three attempts per task, every run priced from one fixed API price list dated 1 September.
Their headline: "The same model can achieve similar success rates at up to 5x costs."
Same model, same tasks, similar results. The bill moved fivefold.
What the spread is made of
Nothing in that range is intelligence. The weights were identical across harnesses. What differed is how much context the harness loads and re-sends on every turn.
The authors give a worked example. Their strongest model solved 97.8% of attempts in Claude Code, and 96.7% in both Codex and Pi. Claude Code cost $1.33 per attempt against Pi's $0.67. Across the shared models on SWE-bench Lite, Claude Code cost about 2.0 times as much as Pi and 1.6 times as much as Codex. Its mean initial context was over ten times Pi's. In nine of twelve comparisons, an alternative harness posted the highest observed success rate.
Read the limits before you use any of that. The authors say their results "can be limited to the two open-source benchmarks we test, which the models may have encountered during training." They also disclose that Arena sponsored the API access for the experiments and runs a leaderboard and coding platform of its own, so it has an interest in the topic. The method is published and the price list is fixed, which is what makes the study usable. It's still a coding benchmark, and your workload isn't one.
The point survives the caveat anyway. The cost of your AI work is set by a pairing decision that nobody wrote down.
The same finding, from the other end
Uber published its own account on 27 August. Between February and mid-August, weekly agent requests grew 9.4 times and weekly active users seven times, while total AI spend, in Uber's words, relatively stabilised from April onward. Holding one model fixed to isolate its own contribution, Uber reports cost per thousand model requests down almost 34% from peak, and cost per session down 52% from its June peak.
None of that came from a discount. Uber broke spend into six multiplied terms (users, sessions per user, turns per session, requests per turn, tokens per request, price per token) and went after the middle three, which is the work the agent does on its own behalf on top of what a person actually asked for.
The levers are unglamorous and specific. Subagents default to a weaker, cheaper model, because a well-specified subtask doesn't need frontier reasoning; Uber calls this the most impactful single default it changed. Prompt cache lifetimes moved from five minutes to an hour, because engineers leave sessions idle and every expired cache rebuilds the prefix at full price. Tool definitions moved out of the session and behind a command-line resolver, removing 50,000 to 70,000 tokens of schema that used to be re-sent on every turn. Batching repeated tool calls into a single script cut tokens on simple warehouse queries by more than half.
These are Uber's figures about Uber, unaudited, and Uber says so: the specific reductions are a product of its codebase, its team and its workflows. The method survives the translation even where the numbers don't.
Most buyers can't see any of this
Ramp's engineering write-up on the cost layer around its internal coding agent describes the ordinary condition. A single flat charge, sitting under one team's budget, turned out to be several separate jobs spanning three repositories, with most of the money spent in follow-on sessions the previous dashboard never counted. Ramp is publishing the fix to a problem its own agent created, which makes it marketing as well as evidence. It's still a fair description of where a typical AI bill goes, which is: nobody is sure.
That's the position most buyers occupy. The invoice arrives as one line from a model provider. The four or five decisions that produced it were made inside an engagement, in week two, usually by accepting whatever the defaults were.
The test
Ask any prospective partner what a comparable piece of work cost them to run last quarter. Not the model's price per million tokens, which is published and which they don't control. The delivered cost of one completed unit of work, off their own bill.
Then ask three follow-ups. Which model and harness handle which step, and whether they tested the alternatives on your kind of work. What the cache lifetime is set to, and why that number rather than the default. What stops a run that isn't converging.
The Berkeley authors recommend comparing the same model's cost and task success across commonly used harnesses, and treating harness complexity as an empirical trade-off. That's a test a partner can run on your workload in a week. A partner who has done this work answers in specifics, including the settings they got wrong first. A partner who hasn't will talk about the model, because the model is the only part of the bill they've ever had to look at.
If you want that question answered for your specific situation, the Forge Playbook does it. Answer a few questions about your business and we'll put together a tailored outline of which workflows are worth automating and what a realistic budget looks like for each. Free, no obligation, takes about three minutes.