The Core Question
How do you know if your engineering team is actually getting better? Not busier. Not more active. Better.
The Measurement Problem
Every engineering leader faces the same challenge: proving that their team is improving. Boards want numbers. Investors want trends. But the numbers most tools provide—commits, lines of code, hours worked—measure activity, not impact.
We spent months studying what actually predicts engineering success. We analyzed research from DORA, SPACE, and academic studies. We talked to CTOs, engineering managers, and individual contributors. We looked at what metrics get gamed, which ones correlate with real outcomes, and why most measurement systems fail.
The result is a framework built around six dimensions of effectiveness. Each dimension answers a specific question about engineering performance. Together, they paint a complete picture that's almost impossible to game.
Why Six Dimensions?
One metric is easy to game. Two metrics create a tradeoff you can exploit. But six interconnected dimensions? Gaming one typically hurts another.
This isn't accidental. It's the core design principle.
Consider the tension: If you optimize purely for delivery speed, quality suffers. If you focus only on quality, delivery slows. If you maximize individual output, collaboration drops. If you spend all your time reviewing others' code, your own delivery tanks.
Effective developers navigate these tradeoffs. The six dimensions capture how well someone balances competing priorities while still shipping meaningful work.
What Are the 6 Dimensions of Developer Effectiveness?
The six dimensions are Delivery (does work ship?), Flow (is effort reaching the finish line sustainably?), Quality (does the work create durable value?), Collaboration (does your presence amplify the team?), Ownership (do you take responsibility for meaningful areas?), and Adaptability (are you improving?). Each answers a specific question about engineering performance. Here's what each one measures and why.
1. Delivery: Does Work Actually Ship?
The Question: Are you completing what you commit to?
Why It Matters: At the end of the day, engineering exists to ship. Strategy documents, architecture discussions, and planning meetings are valuable—but only if they lead to working software in users' hands.
Delivery isn't just about volume. It's about reliability. Can your team predict what they'll accomplish in a sprint? Do completed features stay completed, or do they come back as bugs and rework?
What We Measure:
- Completion Rate (40%): Issues completed vs. assigned. Simple, but foundational.
- Predictability (25%): How consistent is velocity across sprints? High variance suggests estimation problems or scope creep.
- Low Rework (25%): Reverted commits and hotfixes as a percentage of work. Shipping fast only to ship fixes faster isn't progress.
- Estimate Accuracy (10%): How close are actual timelines to estimates? Consistently underestimating (or overestimating) signals planning problems.
The Anti-Gaming Design: You can't just accept fewer issues to boost completion rate—your comparison is against what you committed to. You can't ship broken code faster—rework catches up. You can't pad estimates—accuracy measures deviation in both directions.
2. Flow: Sustainable Efficiency
The Question: Is your cognitive effort reaching the finish line sustainably?
Why It Matters: Context switching destroys developer productivity. Research shows it takes 23 minutes to recover from a single interruption. Developers who start many things but finish few are hemorrhaging cognitive resources. And heroic efforts—80-hour weeks followed by burnout—don't help anyone.
Flow measures the efficiency AND sustainability of your work process. A developer who takes four things from start to finish with consistent output creates more value than one who touches twenty things in bursts and crashes afterward.
What We Measure:
- Cycle Time (25%): How long from starting work to completing it? Faster cycle times mean less work-in-progress inventory.
- WIP Control (25%): The ratio of work assigned to work completed. A ratio of 1:1 is ideal. A ratio of 5:1 means you're juggling too much.
- Batch Size (15%): PR size in lines changed. Too small (under 50 lines) means over-fragmentation. Too large (over 800 lines) means review burden and integration risk.
- Completion Focus (15%): Do you finish things before starting new ones? Starting new work while old work sits incomplete is a flow killer.
- Output Consistency (20%): Consistency of output across sprints. Boom-bust patterns (huge sprint, then barely anything) suggest unsustainable work styles.
The Anti-Gaming Design: You can't game this by submitting tiny PRs (batch size penalty) or huge ones (also penalized). You can't game it by starting lots of work (WIP suffers). You can't hide behind bursts of activity—consistency catches erratic patterns. The only way to score well is to maintain sustainable flow.
3. Quality: Does Your Productivity Create Durable Value?
The Question: Does your code survive contact with reality?
Why It Matters: High throughput with poor quality isn't productivity—it's technical debt accumulation disguised as progress. A developer who ships 50 features that each require 3 bug fixes hasn't shipped 50 features. They've shipped 50 sources of ongoing maintenance.
Quality measures whether your contributions create lasting value or create more work for future you (and future teammates).
What We Measure:
- Defect Density (35%): Bugs introduced relative to work completed. Every feature doesn't need to be bug-free, but patterns matter.
- Stability (30%): How often do your commits get reverted? Reverts are a strong signal that something shipped before it was ready.
- Bug Fix Ratio (20%): Net contribution to codebase quality. Fixed more bugs than you introduced? Bonus. Introduced more than you fixed? That's a concern.
- Incident Avoidance (15%): Hotfixes as a percentage of merged PRs. Hotfixes mean something made it to production that shouldn't have.
The Anti-Gaming Design: You can't avoid bugs by avoiding code—the ratio catches that. You can't hide quality issues by fixing them quickly—stability measures reverts. The only winning strategy is writing quality code in the first place.
4. Collaboration: Do You Make Your Team Better?
The Question: Does your presence amplify team output?
Why It Matters: The best developers aren't just productive individually—they're force multipliers. They review code thoughtfully. They unblock teammates. They share knowledge. A team of collaborators outperforms a team of individual stars every time.
Collaboration measures how much your work helps others succeed, not just how much you personally produce.
What We Measure:
- Review Volume (35%): Reviews given vs. received. Giving more reviews than you receive means you're contributing to team flow.
- Review Responsiveness (25%): How quickly do you review others' code? Long review times are a major source of team friction.
- Unblocking Impact (25%): What fraction of others' PRs do you review? Are you helping keep the team moving?
- Team Contribution (15%): Combined reviews and bug fixes relative to team expectations. Are you pulling your weight on shared responsibilities?
The Anti-Gaming Design: You can't game this by rubber-stamping reviews—quality matters (captured in quality dimension). You can't ignore reviews entirely—volume catches that. The only winning strategy is genuinely helping your team.
5. Ownership: Do You Take Responsibility for Meaningful Areas?
The Question: Do you own outcomes, not just tasks?
Why It Matters: True ownership means caring about the long-term health of your code, not just getting tickets to done. It means doing maintenance work even when it's not glamorous. It means taking on complex problems, not just cherry-picking easy wins.
Ownership measures depth of responsibility—whether you're a tourist passing through codebases or a resident who cares about the neighborhood.
What We Measure:
- Code Area Depth (30%): Consistency of contribution patterns. Do you develop expertise in specific areas, or scatter shallow contributions everywhere?
- Maintenance Investment (25%): Bug fixes as percentage of total work. 10-30% is healthy—it shows you care about code health. 0% suggests you're avoiding tech debt. 50%+ suggests you're only doing reactive work.
- Completion Ownership (25%): Following through on what you start. Starting 10 things and finishing 5 is worse than starting 6 and finishing 6.
- Impact Scope (20%): Complexity of work tackled. Story points per issue vs. team average. Are you taking on meaningful work or just easy wins?
The Anti-Gaming Design: You can't game this by avoiding maintenance (0% maintenance scores 60). You can't game it by only doing bug fixes (low impact scope). You have to actually own areas of the codebase.
6. Adaptability: Are You Getting Better?
The Question: Is your trajectory positive?
Why It Matters: A developer improving from D-level to C-level is more valuable than one stuck at B-level. Growth matters more than static performance. Teams that improve outcompete teams that don't, regardless of starting point.
Adaptability measures the derivative—not where you are, but which direction you're heading.
What We Measure:
- Improvement Rate (30%): Velocity growth over time via linear regression. +10% per sprint is excellent. Flat is concerning. Negative is a problem.
- Quality Improvement (25%): Defect rate trend. Are you introducing fewer bugs over time? Learning from mistakes?
- Efficiency Gains (25%): Cycle time trend. Are you getting faster at completing work? Finding better processes?
- Resilience (20%): Recovery from setbacks. Everyone has bad sprints. How quickly do you bounce back?
The Anti-Gaming Design: This dimension requires at least 2 sprints of data—you can't fake a trend. Improvement has to be real and sustained. One good sprint doesn't move the needle.
The Benefits of Measuring Effectiveness
For Engineering Leaders
Defensible metrics for the board. "Our team's delivery predictability improved 15% quarter-over-quarter while maintaining quality scores above 80" is a statement backed by data that's hard to dismiss.
Early warning system. Declining focus or collaboration scores surface problems before they become crises. You can address burnout risk before losing key people.
Objective performance conversations. Instead of vague feedback, you can point to specific dimensions. "Your delivery is excellent, but your collaboration score suggests you might review more code" is actionable.
For Developers
Clear expectations. The dimensions define what "good" looks like. No more guessing what your manager values.
Growth roadmap. Low score in a dimension? You know exactly what to work on. High score? You know your strengths.
Privacy-first design. Your individual scores are yours. No leaderboards. No comparisons with teammates. Coaching without surveillance.
For Teams
Balanced optimization. When everyone understands all six dimensions, the team naturally balances tradeoffs. No more optimizing delivery at the expense of quality.
Shared vocabulary. "We need to improve our flow" means something specific. Team retrospectives can focus on concrete dimensions.
Culture reinforcement. Measuring collaboration and ownership explicitly signals that these matter—not just shipping features.
Why These Dimensions Work
The six dimensions succeed where other measurement systems fail because they were designed around a single principle: the only way to score well is to actually be effective.
- They measure outcomes, not activity
- They're interconnected, so gaming one hurts others
- They focus on trends, not snapshots
- They preserve privacy while enabling coaching
- They work regardless of what tools developers use
Every dimension answers a real question about engineering effectiveness. Together, they provide a complete picture that no single metric could capture.
The Bottom Line
You can't game six interconnected dimensions. The only winning strategy is to actually be effective.
Getting Started
Measuring developer effectiveness doesn't require new tools or processes. It starts with asking better questions:
- Are we shipping reliably? (Delivery)
- Is work flowing smoothly and sustainably? (Flow)
- Is our output durable? (Quality)
- Are we helping each other? (Collaboration)
- Do we own outcomes? (Ownership)
- Are we improving? (Adaptability)
If you can answer these questions—with data—you understand your team's effectiveness. If you can track them over time, you can prove improvement.
That's the goal. Not surveillance. Not productivity theater. Real measurement of what actually matters.
Measure developer effectiveness, not just productivity
Six dimensions of effectiveness. Trends over time. Insights that help your team see what's working.
Continue reading
- DORA Metrics Without the Dashboard TaxYour CI/CD pipeline already knows how your team operates. We just listen. Why DORA metrics should emerge from integrations you already connected — not another vendor. · 8 min read
- Ship and Stick: How to Measure Whether AI Is Actually WorkingEvery organization is adopting AI. Almost none can prove it's working. Here's how to measure what actually matters — outcomes that ship and stick, not speed that breaks everything. · 12 min read
- The Seven Deadly Sins of Engineering MetricsA field guide to the most toxic measurement patterns in software organizations—and how to avoid them. · 11 min read