Here’s a number that should have ended a lot of dashboards, and didn’t.
When researchers actually measured it — a controlled study of experienced open-source developers, published mid-2025 — AI tools slowed the developers down by 19%. The kicker isn’t the slowdown. It’s that the same developers predicted a 24% speedup going in, and even after they’d finished the tasks, still believed AI had sped them up by 20%. Measured reality and felt reality were pointing in opposite directions (METR / Becker et al., 2025). A room full of smart people, certain they were faster, quantifiably slower — and no instrument on the wall told them.
That gap is the whole post. Because if the people doing the work can’t feel the truth, the dashboard is the only thing that can catch it — and most engineering dashboards are still calibrated for a job nobody does anymore.
The instruments measured the bottleneck. The bottleneck moved.
Think about what velocity, story points, tickets-closed, and lines-of-code actually measured, back when they meant something. All four are proxies for the same underlying scarce resource: human execution. How fast can a person turn a decision into working code. That was the constraint the whole apparatus of scrum was built to see and squeeze. Points-per-sprint went up when your people got faster or your process got less stupid. The proxy tracked the thing.
Everyday — speed in a direction — how fast you're actually moving, and toward what.
In agentic development — story points shipped per sprint. Once a decent proxy for progress; now a vanity metric — it counts execution, the part that just got cheap, and says nothing about whether the right, verified thing got built.
Notice the everyday sense smuggles in a word the scrum sense quietly dropped: direction. Velocity in physics is a vector — speed and heading. Somewhere between the physics classroom and the sprint board we kept the speed and threw away the heading, and nobody flagged it because for twenty years heading and speed were roughly the same problem. When execution is the bottleneck, moving fast and moving toward the goal look almost identical. You couldn’t ship much wrong stuff quickly, because you couldn’t ship much of anything quickly.
That constraint is gone. The models execute now — more of it, faster, than any team you’ll ever staff. Which means the bottleneck slid downstream, off of “can we produce it” and onto “is what we produced actually right, and can we integrate it without setting the codebase on fire.” The old instruments are all still bolted to the wall, still lit up, still measuring the one station on the line that stopped being the constraint.
An instrument pointed at the cheap part is worse than no instrument
You’d think a metric that’s merely gone irrelevant would be harmless — a gauge you learn to ignore. It’s worse than that, and the reason has a name.
Goodhart’s law: when a measure becomes a target, it stops being a good measure. People optimize the number, not the thing the number was standing in for. In the human-execution era, gaming velocity was self-limiting — juicing your points meant a person doing visibly pointless work, and people get bored and caught. The friction was the safety rail.
Agents don’t get bored. Point a coding agent at a velocity target and you can 10x the number this week — more PRs, more tickets, more lines — while the actual health of the system rots underneath, at machine speed. I called this in the series opener: you’re not accumulating tech debt anymore, you’re shipping it at a rate no human team could have produced on its own, and the burndown chart has never looked prettier while the foundation quietly gives way. That’s not a hypothetical. Faros AI, analyzing two years of telemetry across roughly 22,000 developers and 4,000 teams, found engineering throughput up while bugs, incidents, and rework are rising faster — and the constraint quietly relocating to code review, deployment, and QA (Faros AI Impact Report, 2026). The velocity chart is smiling at you from on top of a growing pile of downstream cleanup.
This is the trap of an easy-bake-oven developer with a velocity target: the oven makes output effortless, so output is exactly the wrong thing to reward. A metric that rewards the cheap part doesn’t just fail to help. It hands your fastest gun a loaded incentive to make the mess faster.
What the frontier is actually measuring
So walk the instrument panel over one station, to where the constraint actually lives now. The orgs that are furthest along on this all converge on the same move: stop counting output, start measuring verified outcomes and the new bottlenecks — review, integration, stability, cost. Here’s what that looks like in practice, and who’s saying it.
Pair every speed number with a quality number — never publish one alone. This is the load-bearing rule, and it’s the one DX (getdx.com) is most emphatic about in their AI-measurement guidance: pair throughput with quality “to ensure gains in output aren’t offset by declines elsewhere.” Their canonical quality pair is the PR revert rate — reverted pull requests divided by total pull requests — because a revert is the codebase telling you, in its own hand, that the fast thing was the wrong thing (DX, “5 metrics to measure AI impact”). A speed metric with no quality metric beside it isn’t a measurement. It’s a press release.
Rework and change-failure, not raw commits. The signal isn’t how much code got written — it’s how much of it survived contact with reality. What fraction of AI-authored changes gets reverted, hotfixed, or rewritten within the month. This is the direct read on the failure mode Goodhart guarantees: it catches the agent that ships ten plausible-but-wrong features while a human ships one right one. Faros builds its whole thesis on this gap — individual output up, organizational delivery flat or worse — precisely because throughput and rework moved in opposite directions (Faros AI Impact Report).
Cycle time and review latency — end to end, not per-station. If code generation got 10x faster and lead-time-to-production didn’t move, you didn’t get faster; you just relocated the traffic jam. The frontier watches the whole pipe — commit to verified-in-production — because that’s the only view that exposes the new bottleneck instead of hiding it. Google’s own writeup of the 2025 DORA report lands here hard: AI is “an amplifier.” In cohesive orgs it boosts efficiency; in fragmented ones it magnifies the weakness — and the weakness now is almost always review and integration, not typing speed (Google / DORA 2025). DORA found AI now correlates with higher delivery throughput — a reversal from the prior year — which only sharpens the point: throughput is table stakes; whether it survives is the question.
Cost per verified outcome. This one is genuinely new, and the old scrum vocabulary has no word for it. When execution is metered — tokens, inference, compute — two engineers who “closed the same story” can differ 100x in what it cost to get there. The series opener made the case: the person who ships ten features at ten dollars of inference each is worth more than the one who ships fifteen at three hundred. Volume can’t see that. Cost-per-shipped-outcome is the metric that finally makes leverage legible — and it’s the one your burndown chart is structurally incapable of showing you.
One honest caveat, because Goodhart doesn’t spare the new metrics either. Chase revert rate alone and someone stops writing risky-but-necessary code; chase cost-per-outcome alone and someone ships slop that’s cheap to generate and expensive to live with. That’s the point of the pairing: you’re not swapping one number for a better number, you’re swapping a single vanity number for a small balanced set — speed against quality, output against cost — none of which can be gamed without the paired metric ratting you out. The old dashboard was one gauge because the job had one constraint. The new one has several because the job does too.
The principle underneath all of it
Strip the specific metrics away and what’s left is one rule: measure the verified outcome and the new bottleneck, not the output.
The old stack measured output because output was the bottleneck, so the proxy was honest. Now output is the cheap, abundant, automatable part — and the review is the work now. If your instruments still point at production volume, they’re not neutral-but-outdated. They’re actively lying to you, pointing at the one station where nothing is scarce anymore, and — worse — handing your agents a target that rewards shipping the mess faster.
And this compounds past the individual. As I argued in the team dimension, the unit of work isn’t one dev at a keyboard anymore; it’s a person conducting a fleet of agents, and the whole team conducting fleets in parallel. A velocity chart summed across that team is measurement theater at organizational scale — it aggregates a number that was already lying at the individual level, and prints the total with a straight face.
Go re-read the METR study with that in mind. Those developers felt 20% faster and were 19% slower, and the only thing that could have told them the truth was a measurement they didn’t have. That’s the entire argument, compressed: in an agentic world, your gut is calibrated to the old bottleneck, and so is your dashboard. Both feel right. Both point the wrong way.
So the question I’d actually put to any engineering leader right now: if you overlaid your current sprint metrics on that study — velocity up, tickets up, everyone certain they’re crushing it — would a single number on your board have caught the 19%? Because if the honest answer is no, you’re not measuring an agentic org. You’re taking the temperature of a world that ended, with a thermometer that still reads warm.
This is Part 4 of Leadership in the Agentic Era. Next: The Incongruent Org — what breaks when half your team moves up the stack and half doesn’t.
Comments