Engineering Metrics in the Age of AI

June 21, 2026

Sometime in your first year, a board member will ask whether AI made your engineering team more productive. You will have about ninety seconds and one slide.

This article is about what belongs on that slide.

The trap is that you will not lack numbers. You will have too many, most of them flattering, and every one of them answering a question nobody asked.


Why the Easy Answers Fail#

In 1982, Apple asked Bill Atkinson to report how many lines of code he wrote each week. He had just rewritten a core graphics engine to be faster, smaller, and simpler - so he wrote “-2000.” They stopped sending him the form1.

Engineering has known its output numbers were broken for forty years. We kept using them anyway, for one quiet reason: producing output was expensive.

That cost kept the numbers accidentally honest. Counting pull requests worked the way counting receipts works - forging them took effort. The old rule that a measure stops working once it becomes a target2 stayed half-tamed, because gaming your commit count still meant writing the code.

AI removed the effort. An engineer with an agent can multiply their commits, pull requests, and story points in a week without the product improving at all.

This is happening at platform scale, not just inside your company. In June, GitHub shipped a cap on how many pull requests an outside contributor can have open at once, explaining that maintainers “are dealing with an ever-growing volume of pull requests, including repeated low-quality or drive-by contributions that can slow triage and overwhelm review queues.”3

A volume number that platforms now have to cap is not a volume number that can stand for value.


Diagnostic, Not Dead#

Here is where most of the commentary overshoots, and where I overshot in an earlier version of this piece.

The tempting conclusion is that these numbers are now worthless and should be banned. That is wrong, and a board member who has run an organization will spot it immediately. Volume numbers still carry real information. They have simply lost the one job they were never good at.

The distinction that matters is not good metric versus bad metric. It is diagnostic versus performance measure - what a number can tell you when you are investigating, against what it can prove when you are judging.

Metric What it can legitimately tell you What it cannot tell you
AI usage and adoption Whether the rollout landed, and where it did not Whether AI made anyone more effective
Pull-request count Where work is batching up, and where review is starving How much any individual contributed
Suggestion acceptance rate How well the tool is tuned to your codebase Whether the suggestions were any good
Lines changed The review load and blast radius of a change Value delivered
Code written by AI Roughly how the work is now being produced Anything about quality, speed, or outcome

Every one of those left-hand numbers is useful on a Tuesday when you are trying to work out why a team is stuck.

None of them belongs on the board slide.

The failure mode is not measurement. It is promotion - taking a number that was doing honest diagnostic work and asking it to stand for productivity, which it was never able to do and can now be inflated at will.

That reframing matters practically. “Ban this metric” loses the argument with a CEO who wants a number. “This is a diagnostic, and here is what I use instead when the question is performance” wins it.


What the Honest Numbers Show#

Strip away the keynotes and the measured picture is narrower than either the enthusiasts or the skeptics want.

The industry’s longest-running delivery research has now covered AI adoption in two consecutive annual reports. Between them, the throughput picture improved. The stability picture did not45. Inside the codebase, the two health indicators move in opposite directions: duplication up, refactoring down6.

And the most useful benchmark I have found puts the median gain in pull-request throughput at 7.8%, with most organizations landing between 5% and 15%7. This is vendor research rather than peer review, drawn from more than 400 companies, and it is worth noting that it measures throughput of pull requests - itself a generation-side proxy. Even the friendly number, measured on friendly terms, is single digits.

Genuinely useful. An order of magnitude below what is claimed on stage.

Two of the three claims currently circulating in board decks fail on definition alone. Microsoft and Google have both cited figures around 30% of code written by AI, without a shared definition of what counts8 - it is the lines-of-code metric reborn with better marketing. Suggestion acceptance rate measures how agreeable your engineers are. And near-universal AI usage5 is table stakes, not a result.

If your dashboard says AI tripled your productivity, your dashboard is measuring generation rather than delivery. Why those two diverge is its own article9.


The Paired Scorecard#

The numbers that survive share one property: they describe the system, not a person. No amount of individual volume can fake them. As Dan North puts it, they “measure the engine, not the contribution of individual pistons.”10

But system metrics alone are still gameable in one direction - you can always buy speed with quality. So the working rule is: never put a speed number on the slide without the counter-metric that catches what it costs.

Outcome Speed signal Counter-metric What triggers a decision
Delivery Lead time from start to users Change-failure rate Both rising: you bought speed with stability. Investigate before celebrating.
Throughput Merged changes Rework within 14 days Rework climbing: reduce batch size before adding capacity.
Maintainability Share of effort on refactoring Duplication growth Duplication up while refactoring flat: reserve capacity now, or pay in hesitancy later11.
Developer experience Self-reported flow time Interruptions and handoffs Flow falling while throughput holds: you are burning goodwill to hit numbers.

Three rules of hygiene make it work. Measure teams and services rather than individuals. Watch trends rather than levels - you are looking for a direction, not a grade. And give every number an owner and a threshold, because a metric that cannot trigger a decision is decoration12.


The Ninety Seconds#

Here is the scorecard doing its job in an actual sentence:

“Delivery lead time improved 14% this half, and our change-failure rate stayed flat. We read that as a real gain rather than volume moving downstream - if the speed were coming out of quality, failures would have risen with it. Rework is up slightly, so we are cutting batch size next quarter. We are not reporting an AI productivity number, because the ones available measure how much the tools are used rather than what shipped.”

That answer takes twenty seconds, survives a follow-up question, and does something the “AI made us 40% faster” answer cannot: it tells the board what you will do next.

The last sentence is the one that takes nerve. Say it anyway. Declining to report a bad number is a stronger position than reporting one you will have to walk back when the quarter does not match it.


Where I Would Push Back on Myself#

The strongest objection to everything above is that “never measure individuals” is too convenient.

Managers do have to evaluate individuals. Promotions, performance reviews, and the occasional decision to let someone go are real and unavoidable. Refusing to measure people does not make those judgments disappear. It just moves them somewhere unaccountable - into visibility, likeability, and how recently someone spoke in a meeting. That is not obviously fairer than a flawed number, and for people who are quiet or remote or new, it is usually worse.

So I would state the rule more narrowly than I used to.

Individual assessment is legitimate. Individual metrics are the problem, and specifically metrics derived from output volume, because those are exactly the ones AI can now inflate on demand. What survives at the individual level is evidence a person’s peers would recognize: what they own, what they unblocked, whether their reviews improve the work, whether their judgment is trusted on the thing they are responsible for.

That is harder than a dashboard. It is also the part of management that was never going to be automated.


Closing#

AI did not break engineering metrics. It broke the illusion that the broken ones were working - teams already measuring outcomes barely had to change anything, and everyone else lost their alibi.

Atkinson’s weekly form had it backwards in 1982 and it is still backwards now. Some of the best engineering work of the next few years will look like negative two thousand lines, and your scorecard should be able to tell.


References#


  1. “-2000 Lines of Code,” Folklore.org, Andy Hertzfeld’s Macintosh stories. https://www.folklore.org/Negative_2000_Lines_Of_Code.html ↩︎

  2. Marilyn Strathern, “‘Improving Ratings’: Audit in the British University System,” European Review (1997), crediting Keith Hoskin; C.A.E. Goodhart, “Problems of Monetary Management: The UK Experience” (1975). ↩︎

  3. GitHub Changelog, “Limit open pull requests for users without write access” (June 2026). https://github.blog/changelog/2026-06-17-limit-open-pull-requests-for-users-without-write-access/ ↩︎

  4. DORA, Accelerate State of DevOps Report 2024. https://dora.dev/research/2024/dora-report/ ↩︎

  5. DORA, State of AI-assisted Software Development 2025; around 90% of developers report daily AI use. https://dora.dev/dora-report-2025/ ↩︎ ↩︎

  6. GitClear, AI Copilot Code Quality research (2025). https://www.gitclear.com/ai_assistant_code_quality_2025_research ↩︎

  7. DX, “The AI Efficiency Plateau” (May 2026), analyzing pull-request throughput across more than 400 companies. https://getdx.com/blog/the-ai-efficiency-plateau/ ↩︎

  8. TechCrunch, “Microsoft CEO says up to 30% of the company’s code was written by AI” (April 2025). https://techcrunch.com/2025/04/29/microsoft-ceo-says-up-to-30-of-the-companys-code-was-written-by-ai/ ↩︎

  9. Code Was Never the Bottleneck. https://avivzaken.com/docs/code-was-never-the-bottleneck/ ↩︎

  10. Dan North, “The Worst Programmer I Know” (2023). https://dannorth.net/blog/the-worst-programmer/ ↩︎

  11. Technical Debt Is a Tool. https://avivzaken.com/docs/technical-debt-is-a-tool/ ↩︎

  12. Staying One Step Ahead. https://avivzaken.com/docs/staying-one-step-ahead/ ↩︎

Previous Evals Are the New Tests Next The Shape of the Team Is Changing

Aviv Zaken

Founder · CTO · Engineering Leader · Entrepreneur