Blog

220% More Code, 36% More Product

Reuters reported four numbers from inside Meta's Project OT: code changes up 220%, features reaching users up 36%, incidents up 40%, firefighting time up 70%. The company with the least possible access constraint still couldn't take delivery of its own output. Three numbers to put next to each other before you conclude AI is working.

FForge
··5 min read
Listen to this article 0:00 / 6:40
  • AI implementation
  • implementation reality
  • measurement
  • Meta
  • McKinsey
  • agentic development
  • productivity
220% More Code, 36% More Product
Photo by Alberto Rodríguez on Unsplash

Reuters published an investigation on August 26 into Project OT, Meta's plan to restructure itself around AI agents. The part worth keeping isn't the part that got quoted. Sitting inside the reporting are four numbers from Meta's own internal posts, and read in the order the work actually flows, they describe a pipeline rather than a scandal.

According to Reuters, a June post from Meta's chief technology officer put code changes to the company's AI platforms and infrastructure up 220% year over year. Changes that produced new or improved features reaching users were up 36%. Technical and security incidents were up 40%. Employee time spent on those incidents was up 70%.

Work created. Work delivered. Things broken. Hours spent fixing them. That's the whole pipe, instrumented at four points, at one of the few companies on earth with the means to instrument it and the nerve to write the results down.


Nothing here was an access problem

Start with what Meta wasn't short of.

It had frontier models, including the ones it built itself. It had the researchers who built them. It had compute that no mid-market buyer will ever be quoted a price for. It had the mandate coming from the CEO, out of a January leadership retreat, which removes the sponsorship gap that quietly ends most internal programs by month four. And it had engineers who are, by any reasonable measure, among the most capable available.

If the constraint on getting value out of AI were access to capability, Meta cleared it more completely than any organization in the world.

The first number says the capability did its job. Code changes more than tripled year over year. Whatever you make of agentic development, something in that stack produced an enormous amount of output.

The next three say the organization couldn't take delivery of it. Output climbed steeply. Features reaching users climbed modestly. Incidents rose, and the hours spent cleaning up after those incidents rose faster than the incidents themselves, which is the signal that the cleanup was getting harder rather than just more frequent.

That's not a model failure. Nothing in those four numbers suggests the agents were bad at writing code. They suggest the agents were good at writing code, and that writing code was never the part of the system with the least capacity.

Meta isn't an outlier, either. McKinsey's Technology Trends Outlook 2026, released in September, found that productivity fell in 30% of companies after their teams started using agentic AI tools. In one case the report cites, coding activity rose 180% while shipped releases rose 30%. Different company, different instrumentation, same shape: the output side of the pipe grew far faster than the delivery side.


Why one number is easy and three are not

Every organization running AI right now can produce the first number. Volume of output is the cheapest thing in the pipeline to measure, because the tools emit it by default: commits, tickets closed, documents drafted, calls summarized. It arrives on a dashboard without anyone deciding to build one.

The other three take work. Throughput to the customer means agreeing what counts as delivered, which is a conversation about definitions most teams have never had to hold. Defect rate means someone attributing an incident back to the thing that caused it. Remediation time means tracking hours against cleanup rather than against features, which most time-tracking is actively designed not to do.

So the measurement that's easy to collect is the one that can move a great deal while the business gets nothing, and the measurements that would catch that are the ones nobody has. When a vendor or an internal team reports an AI productivity gain, ask which of the four they measured. In practice it's almost always the first, because it's the only one available by default, not because anyone decided it was the right one.

This is the part that generalizes down from Meta's scale. The budget doesn't generalize and neither does the engineering bench. The instrumentation gap does, and it gets worse further down market, because a smaller organization has fewer places for surplus work to sit visibly before it turns into an incident.


What Reuters did and didn't establish

Meta's position, reported alongside the rest, is that the more aggressive headcount reductions were scenarios under consideration rather than decisions. That's worth carrying at full strength rather than as a footnote. Reuters also reports the broader plan hasn't been abandoned: a November wave was scrapped, not the strategy.

And Reuters did not establish why that wave was canceled. Weak results plus internal backlash is the obvious reading, and it's the one most coverage reached, but it's a reading rather than a finding. Treat it as inference, including this one.

What does stand up is narrower and more useful. At a company with no capability constraint, output rose far more than delivery, and the difference showed up as incidents and firefighting hours.


The test

Take your last quarter of AI-assisted work, in whatever function has gone furthest, and put three numbers next to each other.

How much more got produced. How much more actually reached a customer, closed a case, or shipped. How many hours went into fixing what came back.

If you can only produce the first number, that's the finding. You don't have a measurement problem you'll get to later; you have a decision you're currently making without the two numbers that would tell you whether it's working. Meta could produce all three and chose to look at them, which is why its program got stopped in month five rather than month twenty.

Most organizations can't, and won't find out for considerably longer than that.


If you want that question answered for your specific situation, the Forge Playbook does it. Answer a few questions about your business and we'll put together a tailored outline of which workflows are worth automating and what a realistic budget looks like for each. Free, no obligation, takes about three minutes.

Get your free Forge Playbook →

Ashton & ForgeAshton & Forge

We vet the agencies, match you with the right three, and give you the plan to brief them.

/Subscribe to Updates

The occasional brief. No spam, unsubscribe anytime.

© 2026 Ashton & Forge