In this one-off interactive, gamified workshop, we’ll simulate real-world work scenarios at your organisation via a board game, helping you identify and eliminate bottlenecks, inefficient processes, and unhelpful feedback loops.
Workshop Details
Teams from startups to enterprises are writing far more code, roughly 17 times more according to the NBER working paper Laura cites, and change sets have almost doubled in size across several independent datasets. What has not changed much is what reaches production. Very few organisations have scaled outcomes anywhere near as fast as they scaled code generation. Laura's framing is a J curve, with the caveat that the future is unevenly distributed and some organisations are experiencing this very differently from others.
Then the question she gets asked weekly: does AI need new metrics? No. What it needs is the discipline most companies skipped the first time.
"AI just amplifies everything. It amplifies the good, and it amplifies the bad."
When leaders bring her their AI measurement problem, she usually finds they weren't measuring efficiency or performance well before AI either. So the conversation turns to what a serious approach actually looks like: several metrics rather than one number, a mix of self-reported and system data, and metrics held in tension like a spider web, so PRs per engineer gets pulled on by change failure rate, which gets pulled on by developer experience and cycle time. Nobody can over-optimise one dimension without the whole thing collapsing. And every metric has to be tied to a decision somebody actually makes. The moment a number lands on a dashboard it creates an incentive, whether you intended one or not.
Laura is blunt about what does harm. Individual-level measurement is a red flag, which puts token leaderboards, commits per engineer, and PR stack ranking in the same bin. She tells the Shopify story fairly: it started as a sensible way to signal that spending money on tokens was fine, executives included, back when nobody knew whether it was acceptable. It worked. It also needed an end date, and the companies still running them are now watching them incentivise the wrong behaviour.
For anyone at a small or mid-sized company with no measurement practice at all, the practical advice is to stop waiting for a data platform and send a survey. Not a happiness survey. Ask people how many PRs they closed, how confident they are that what they ship won't break. Then look at what the answers tell you to investigate next.
Roughly 17x more code is being written, and only a fraction of that shows up in shipped releases. The gap is where the productivity story falls apart.
AI doesn't require new metrics. It requires the measurement discipline most organisations never built.
Any engineering metric measured at the individual level is a red flag. Token leaderboards had a legitimate window, and for most companies that window has closed.
Metrics should be held in tension with each other, and each one should be tied to a decision someone has to make. A dashboard built out of curiosity still changes behaviour.
Self-reported survey data gets you far enough to act. You do not need a data engineering project first.
"These metrics are not answers. They are questions."
Laura Tacho is a Senior Principal Technologist at AWS working on developer experience and AI-native development. She was previously CTO at DX, an engineering intelligence platform, where she co-authored the DX Core 4 framework for measuring developer productivity, now used at companies including Netflix, Airbnb, and LEGO. She has been building developer tools for 15 years, from the early days of IaaS and PaaS through Docker and CI/CD, and coaches engineering leaders from startups to the Fortune 500. She was named Austrian Innovator of the Year in 2025.
Laura's website, research, and metrics course
Free developer experience survey template
A deep dive into developer experience surveys
Writing Code vs. Shipping Code, NBER Working Paper 35275
DX data on AI-authored code and PR size
Jellyfish AI Engineering Trends
The Pragmatic Engineer on tokenmaxxing and the Shopify leaderboard
WeAreDevelopers World Congress
Find the Stellar Work Podcast on Spotify, Apple Podcasts, YouTube, and more.
For weekly essays on transformation, flow, and AI in knowledge work, join the Stellar Work newsletter.
The Stellar Work Podcast is hosted by Ben, founder of Stellar Work. Conversations with the people shaping how work actually gets done.
Not Sure Where to Start?
In this one-off interactive, gamified workshop, we’ll simulate real-world work scenarios at your organisation via a board game, helping you identify and eliminate bottlenecks, inefficient processes, and unhelpful feedback loops.
Workshop Details