The Measurement Gap
Every organization is adopting AI. Almost none can measure whether it's working. The old metrics are broken. The new ones don't exist yet. This is the gap.
Your organization just rolled out AI coding assistants. The CTO is asking for ROI numbers. The vendor dashboard says adoption is at 70%. Developers report feeling more productive. Everything looks great.
Except nobody can actually prove anything changed.
The DORA 2024 State of DevOps Report found that teams using AI coding assistants saw a 1.5% decrease in throughput and a 7.2% decrease in stability1. Not improvement. Decrease. Meanwhile, a METR study found that experienced developers using AI completed tasks 19% slower — while believing they were 20% faster2.
Read that again. They felt faster. They were slower.
This is not a knock on AI. AI is transformative. But the measurement infrastructure most organizations are using to evaluate AI — and every other change they're making — is fundamentally broken.
"Go Fast and Break Everything" Is Not a Strategy
There's a seductive narrative in software right now: AI makes developers faster, faster is better, therefore AI is better. Ship more. Commit more. Close more tickets.
This is the factory mentality applied to knowledge work. And it creates the same problem factories discovered decades ago: speed without quality control is waste.
Uplevel's 2024 analysis of roughly 800 developers found a 41% higher bug rate among teams using AI coding assistants3. Faros AI reported teams closing 21% more tasks but with 91% longer code reviews and 9% more bugs4. The code ships faster and breaks more.
You wouldn't run a factory without quality control. You wouldn't celebrate a production line that doubled output while tripling defects. Yet that's exactly what happens when engineering organizations measure AI success by throughput alone.
The problem isn't AI. The problem is measuring speed without measuring durability. You're checking the heart rate but ignoring the blood pressure.
The Speed Trap
High velocity plus high rework is not productivity. It's technical debt with a faster accumulation rate.
And organizations are spending real money on this blind spot. AI tool licenses, infrastructure costs, training programs — all evaluated on adoption rates and subjective developer satisfaction surveys. No feedback loop. No outcome measurement. No way to know if the investment is paying off or making things worse.
What Does "Ship and Stick" Actually Mean?
"Ship and stick" is the question that cuts through the noise: did the work ship, and did it stick?
Not "did it ship fast." Not "did it close a ticket." Did it create durable value? Did it survive in production? Did it solve the problem it was supposed to solve without creating three new ones?
"Ship and stick" is a simple concept with concrete, measurable signals. Think of these as the vital signs of your engineering organization's change health — like a doctor checking blood pressure, oxygen levels, and heart rate together, not just one in isolation.
The Vital Signs
| Signal | What It Measures | Why It Matters |
|---|---|---|
| Rework rate | Reverts and hotfixes as a ratio of shipped work | Code that gets reverted is waste, not productivity — regardless of how fast it was written |
| Quality trajectory | Defect density trend over time | One bad sprint is a blip. Declining quality over six sprints is a systemic problem |
| Predictability | Consistency of delivery across sprints | Sustainable improvement is consistent, not boom-bust. A team swinging between 60% and 100% completion is less healthy than one consistently at 80% |
| Sustainability | Declining completion patterns, boom-bust cycles | If output spikes then crashes, the team is sprinting, not running. That pace will break |
These aren't theoretical. They're measurable from the data your team already generates — issues completed, commits merged, PRs reviewed, bugs reported.
The key insight is that no single signal tells the story. A team can have excellent velocity and terrible rework. They can have low bug counts but declining predictability. The vital signs work together, like a diagnostic panel.
A team that ships and sticks looks like this: consistent delivery, stable or improving quality, manageable rework, sustainable pace. That's what real productivity looks like — whether they're using AI or not.
A team that's "going fast and breaking everything" looks like this: high velocity plus high rework. Lots of code merged, lots of it reverted. Sprint output swinging wildly. Bug density creeping up. Speed that creates more work than it eliminates.
The AI-Neutral Principle
Here's a position that makes some people uncomfortable: we don't care what tools you use.
We don't track whether a developer used Copilot, Claude, or a mechanical keyboard and pure willpower. We don't measure AI adoption rates. We don't count AI-generated lines of code.
We measure what ships. We measure what sticks.
The AI-neutral principle is measuring outcomes without tracking which tools produced them. It matters because it's the only honest way to evaluate any change to how your team works.
Think about what "effective AI usage" actually looks like in outcomes:
- Higher throughput with stable or improved quality
- Reduced cycle time without increased rework
- More work shipped that stays shipped
And what "poor AI usage" looks like:
- Speed plus instability — fast commits followed by frequent fixes
- More code produced but more of it reverted
- Shorter time to merge but longer time to stabilize
The METR study's perception gap is the poster child for why you need outcome measurement, not opinion surveys. Developers felt 20% faster. They were 19% slower2. Without outcome data, you'd celebrate the adoption and miss the regression.
The Only Honest Question
Don't ask "Are people using AI?" Ask "Are people effective?" If the outcomes are better, the tools are working. If they're not, the tools aren't — regardless of adoption rates.
This principle extends beyond AI. New processes, team restructures, methodology changes — every organizational change promises improvement. The question is always the same: did outcomes actually improve, or did it just feel like they did?
The organizations that win aren't the ones that adopt the fastest. They're the ones that can prove their changes are working.
The Signals Your Team Is Actually Adapting
Here's something that gets overlooked in the metrics conversation: the "hard" numbers only tell half the story. The other half is the human side — whether the team is genuinely adapting to change or quietly falling apart under it.
Engineering teams navigating change — AI adoption, new processes, reorgs — generate signals that predict whether the change will stick long before the delivery metrics confirm it. We think of these as the vital signs of team health.
Mood Trajectory
Direction matters more than absolute level. A team whose mood is trending upward from a 3 to a 4 is healthier than a team sitting flat at a 7. Flat-high can mean complacency. Rising means momentum.
Declining mood after a major change — say, an AI tool rollout — is an early warning that something isn't landing well. You'll see it in mood data weeks before it shows up in delivery metrics.
Candor in Retrospectives
A team that only posts positive retro cards isn't a happy team. It's a team that doesn't feel safe being honest.
Healthy teams have a balance of positive and improvement-focused feedback. Research on psychological safety consistently shows that the best teams surface problems openly5. A candor ratio that skews too positive — everyone saying things went great when clearly they didn't — is a red flag for suppressed concerns.
Anonymous card rates tell a similar story. If more than half the retro cards are anonymous, people may not feel safe attaching their name to honest feedback. That's a change health problem, not just a process problem.
Action Follow-Through
This is the one that separates teams that learn from teams that just talk. A PMI community poll found that nearly two-thirds of teams implemented fewer than 25% of their retrospective action items6. The pattern: identify problems, commit to fixes, do nothing, repeat.
Action follow-through is the most concrete signal of whether a team is actually adapting. It answers: when the team commits to a change, does the change happen? If you're tracking AI adoption impact and the team keeps identifying integration friction in retros but never resolves it, no amount of tooling will help.
Resilience
Every team has bad sprints. What matters is what happens next.
A team that drops from 85% to 60% completion and bounces back to 80% the following sprint is showing genuine adaptability. A team that drops and stays down is showing that the change overwhelmed their capacity to absorb it.
Resilience — the ability to recover from setbacks — is one of the strongest predictors of whether a team will successfully navigate change over time. Teams that bounce back are learning. Teams that don't are stuck.
Why These 'Soft' Signals Matter
Delivery metrics tell you what happened. Team health signals tell you what's about to happen. A team with strong delivery but declining mood and low action follow-through is a team about to hit a wall. The numbers look fine today. They won't in two months.
Regardless of How — Value Is the Only Metric That Matters
Let's zoom out.
Whether the change is AI tools, a new sprint cadence, a team restructure, or a methodology shift — the question is always the same: is it working?
Not "did we adopt it?" Not "do people like it?" Not "does the vendor dashboard look good?"
Is the team shipping work that sticks? Is quality stable or improving? Is the pace sustainable? Are people actually adapting, or just going through the motions?
This is what we mean by "what ships and sticks matters most, regardless of how." The tools, processes, and structures are inputs. Value creation is the output. If you can't measure the output, you're flying blind — optimizing inputs and hoping for the best.
The Feedback Loop
The organizations that get this right build a continuous feedback loop:
- Make a change — adopt a tool, adjust a process, restructure a team
- Measure the outcome — did delivery improve? Did quality hold? Is the team adapting?
- Adjust based on evidence — double down on what's working, course-correct what isn't
- Repeat — every sprint, every quarter, continuously
This sounds obvious. Almost nobody does it. Most organizations are stuck at step 1 — making changes and assuming they worked because they feel right.
The measurement gap isn't a technology problem. It's an organizational discipline problem. The data exists. Your project management tools, code repositories, and team ceremonies already generate the signals you need. The question is whether you're looking at them — and whether you're looking at the right ones.
Health Scores as the Measurement Layer
We built Simyl Flow around this idea: a single health score that synthesizes delivery, quality, sustainability, and team dynamics into one number that answers "are we getting better?"
Not a vanity metric. Not a surveillance tool. A vital sign — like a patient's chart that tells the doctor whether the treatment is working, without micromanaging which pills the patient took at which hour.
The health score goes up when outcomes improve: more work ships and sticks, quality holds, the team is adapting. It goes down when outcomes degrade: rework increases, predictability drops, the team is showing signs of change fatigue.
It's the feedback loop that most organizations are missing. Not another dashboard of activity metrics. A single, evidence-based answer to the question every engineering leader needs answered: is what we're doing actually working?
The Bottom Line
Every organization is making changes. New tools, new processes, new structures. Almost none of them can prove whether those changes are working.
The teams that win — the ones that genuinely improve rather than just churn — share three characteristics:
- They measure outcomes, not activity. Not commits, not adoption rates, not story points. Did value ship? Did it stick?
- They watch the human signals. Mood trajectory, candor, action follow-through, resilience. The vital signs that predict whether change will hold.
- They build feedback loops. Change, measure, adjust, repeat. Every sprint. Without exception.
Ship and stick. That's the standard. Everything else is noise.
Measure developer effectiveness, not just productivity
Six dimensions of effectiveness. Trends over time. Insights that help your team see what's working.
Sources
Footnotes
-
DORA (2024). Accelerate State of DevOps Report — Teams using AI coding assistants experienced 1.5% decrease in delivery throughput and 7.2% decrease in delivery stability. ↩
-
METR (2025). Measuring the Impact of Early AI Assistance on Software Development — Experienced open-source developers completed tasks 19% slower with AI assistance, while perceiving themselves as 20% faster. ↩ ↩2
-
Uplevel (2024). Can Generative AI Improve Developer Productivity? — Analysis of ~800 developers showing a 41% higher bug rate among teams using GitHub Copilot, with no significant change in cycle time. ↩
-
Faros AI (2024). State of Software Development Report — Teams closing 21% more tasks with AI assistance, but experiencing 91% longer code reviews and 9% more bugs. ↩
-
Edmondson, A. (2018). The Fearless Organization: Creating Psychological Safety in the Workplace for Learning, Innovation, and Growth — Teams with high psychological safety consistently outperform those without it. ↩
-
PMI Community Poll (2022), reported in Bondale, K. "Why hold retrospectives if ideas don't get implemented?" — Nearly two-thirds of respondents reported implementing fewer than 25% of retrospective improvement ideas. ↩
Continue reading
- Did It Stick? And What Did It Cost?Every CFO is asking what the AI bill is. Most engineering leaders can produce an invoice and a vibe. Here's the missing half of AI ROI — and why no one else can show it to you. · 8 min read
- The 6 Dimensions of Developer Effectiveness: A Framework for Measuring What Actually MattersWhy we chose these specific dimensions, what each one reveals about real engineering performance, and how measuring outcomes transforms teams. · 10 min read
- Twelve Working Agreements for Machine-Written CodeThe vibe-coding hangover retro ends with rules on a whiteboard. Here are twelve you can steal — each one a single-sentence rule, the number that moves if it's holding, and the check-in that keeps it honest. · 13 min read