#64 Laura Tacho - You weren't measuring well before AI either

August 24, 2026
\
Ben
\
26:44
 mins

AI did not break engineering metrics. It made it obvious that most companies never had any worth trusting. Laura Tacho has spent 15 years in developer tools, ran DX as CTO, and now works on developer experience at AWS. That path has given her an unusual vantage point: she has seen the actual measurement data from hundreds of engineering organisations, at exactly the moment when every CEO started asking why their team isn't 10x yet. Her answer is less comfortable than either the hype or the backlash.

Teams from startups to enterprises are writing far more code, roughly 17 times more according to the NBER working paper Laura cites, and change sets have almost doubled in size across several independent datasets. What has not changed much is what reaches production. Very few organisations have scaled outcomes anywhere near as fast as they scaled code generation. Laura's framing is a J curve, with the caveat that the future is unevenly distributed and some organisations are experiencing this very differently from others.

Then the question she gets asked weekly: does AI need new metrics? No. What it needs is the discipline most companies skipped the first time.

"AI just amplifies everything. It amplifies the good, and it amplifies the bad."

When leaders bring her their AI measurement problem, she usually finds they weren't measuring efficiency or performance well before AI either. So the conversation turns to what a serious approach actually looks like: several metrics rather than one number, a mix of self-reported and system data, and metrics held in tension like a spider web, so PRs per engineer gets pulled on by change failure rate, which gets pulled on by developer experience and cycle time. Nobody can over-optimise one dimension without the whole thing collapsing. And every metric has to be tied to a decision somebody actually makes. The moment a number lands on a dashboard it creates an incentive, whether you intended one or not.

Laura is blunt about what does harm. Individual-level measurement is a red flag, which puts token leaderboards, commits per engineer, and PR stack ranking in the same bin. She tells the Shopify story fairly: it started as a sensible way to signal that spending money on tokens was fine, executives included, back when nobody knew whether it was acceptable. It worked. It also needed an end date, and the companies still running them are now watching them incentivise the wrong behaviour.

For anyone at a small or mid-sized company with no measurement practice at all, the practical advice is to stop waiting for a data platform and send a survey. Not a happiness survey. Ask people how many PRs they closed, how confident they are that what they ship won't break. Then look at what the answers tell you to investigate next.

Key takeaways

Roughly 17x more code is being written, and only a fraction of that shows up in shipped releases. The gap is where the productivity story falls apart.

AI doesn't require new metrics. It requires the measurement discipline most organisations never built.

Any engineering metric measured at the individual level is a red flag. Token leaderboards had a legitimate window, and for most companies that window has closed.

Metrics should be held in tension with each other, and each one should be tied to a decision someone has to make. A dashboard built out of curiosity still changes behaviour.

Self-reported survey data gets you far enough to act. You do not need a data engineering project first.

"These metrics are not answers. They are questions."

About the guest

Laura Tacho is a Senior Principal Technologist at AWS working on developer experience and AI-native development. She was previously CTO at DX, an engineering intelligence platform, where she co-authored the DX Core 4 framework for measuring developer productivity, now used at companies including Netflix, Airbnb, and LEGO. She has been building developer tools for 15 years, from the early days of IaaS and PaaS through Docker and CI/CD, and coaches engineering leaders from startups to the Fortune 500. She was named Austrian Innovator of the Year in 2025.

Resources mentioned

Laura Tacho on LinkedIn

Laura's website, research, and metrics course

Free developer experience survey template

A deep dive into developer experience surveys

The DX Core 4 framework

Writing Code vs. Shipping Code, NBER Working Paper 35275

DX data on AI-authored code and PR size

Jellyfish AI Engineering Trends

The Pragmatic Engineer on tokenmaxxing and the Shopify leaderboard

WeAreDevelopers World Congress

Listen & subscribe

Find the Stellar Work Podcast on Spotify, Apple Podcasts, YouTube, and more.

For weekly essays on transformation, flow, and AI in knowledge work, join the Stellar Work newsletter.

The Stellar Work Podcast is hosted by Ben, founder of Stellar Work. Conversations with the people shaping how work actually gets done.

Not Sure Where to Start?

Warp Speed Workshop

In this one-off interactive, gamified workshop, we’ll simulate real-world work scenarios at your organisation via a board game, helping you identify and eliminate bottlenecks, inefficient processes, and unhelpful feedback loops.

Close Cookie Popup
Cookie Preferences
By clicking “Accept All”, you agree to the storing of cookies on your device to enhance site navigation, analyze site usage and assist in our marketing efforts as outlined in our privacy policy.
Strictly Necessary (Always Active)
Cookies required to enable basic website functionality.
Cookies helping us understand how this website performs, how visitors interact with the site, and whether there may be technical issues.
Cookies used to deliver advertising that is more relevant to you and your interests.
Cookies allowing the website to remember choices you make (such as your user name, language, or the region you are in).