Two Questions, One Answer
"Is AI working?" is two questions, not one. Did it stick? And what did it cost? Most teams can't answer either. The ones that can are pulling away.
Your CFO walked into your one-on-one last week with a number on a sticky note. It was bigger than they expected. They want to know what you got for it.
Six months ago, AI coding tools were a developer-side curiosity — a few seats here, a side-by-side trial there. They're now a line item. Cursor, Copilot, Claude, Codeium, the model-API budget that quietly tripled — most engineering organizations are spending real money on AI, and most can't tell you whether the spend produced anything that stuck.
This is the second half of Ship and Stick. The first half answered: did the work the AI helped with actually ship and survive? The second half is the question every finance partner is now asking, and most engineering leaders can only answer with an invoice and a vibe:
"What did it cost us, and was it worth it?"
A Reminder: Ship and Stick
We argued earlier this year that activity metrics — adoption rates, prompts written, lines accepted — are vanity. The only honest measure of whether AI is working is whether the outcomes shipped and stuck. Did the PR merge? Did it stay merged? Did quality hold? Did predictability improve, or did the team start swinging?
That's still the right frame. But it answers half a question.
Outcomes alone tell you whether something improved. They don't tell you whether the improvement was worth what you paid for it. A 4% lift in delivery health that costs $80k a quarter in tool licenses is a different bet than a 4% lift that costs $8k. Without the denominator, you can't tell which one you're sitting on.
How Do You Measure AI ROI?
AI ROI is the value of what stuck divided by what you paid. Here's the math every engineering leader is being asked to do:
AI ROI = (value of what stuck) / (what you paid)
Most teams are missing the denominator entirely. They have a vendor invoice in finance's inbox and a sense that "people seem more productive." That's not ROI. That's hope plus a credit card statement.
The denominator has moved. AI spend used to be a footnote — twenty bucks a seat for an editor plugin. In 2026 it's a column:
- Per-seat licenses across multiple coding tools
- API usage charges that scale with how much your team actually uses them — the more it works, the more you pay
- Model-routing services, vector databases, evaluation tooling
- The internal AI features your own product is shipping, which run on the same model providers
Three things follow from that.
One: AI spend has become non-trivial enough that finance is now a stakeholder in engineering decisions, whether you like it or not.
Two: AI spend is no longer fixed-cost. It scales with usage — which means a successful adoption increases the bill. Winning the adoption fight without the outcome instrumentation is how you end up explaining a 3x year-over-year invoice to a CFO who can't see the impact.
Three: The two numbers move on different clocks. Spend is monthly and visible. Outcomes are sprint-over-sprint and invisible without a measurement layer. If you only see one of them, you'll make decisions on the one you see — and the one you see is the bill.
The Asymmetry Trap
Spend is loud. Outcomes are quiet. Without instrumentation, you'll cut tools that were working and keep tools that weren't — because the bill arrives on time and the impact never does.
Why No One Shows You Both Halves
If pairing outcomes with spend is so obviously useful, why hasn't your AI vendor built it for you?
Because the math doesn't work for them.
| Vendor type | What they show you | What they don't |
|---|---|---|
| AI coding tools | Adoption, accepted suggestions, "time saved" (estimated by them) | The trend in your delivery health. Whether your bug rate moved. The bill in context. |
| Productivity dashboards | Cycle time, PR throughput, individual leaderboards | What you paid. Whether the activity produced durable outcomes. |
| FinOps / spend tools | The invoice, broken out by service | What any of it shipped. Whether anything stuck. |
| Engineering surveys | Self-reported satisfaction with AI tools | The perception gap — what people felt versus what they did. |
Every one of those tools is shipping the half they can sell, and quietly omitting the half that would make the picture honest. A coding tool that showed you cost-per-stuck-PR would be marking its own homework. A productivity dashboard that exposed the bill would be giving you a reason to cancel it.
The closed loop — cost paired with outcome, both trending sprint over sprint — isn't a feature any single-purpose vendor can ship without undermining itself.
What the Loop Looks Like
The instrumentation isn't complicated. It's just rare.
You need three things on one screen:
- A delivery health trend. Sprint over sprint, AI-neutral. Did work ship? Did it stick? Are quality and predictability holding?
- An AI spend trend. Same cadence. Per team, not per developer. (More on that in a moment.)
- A causal hypothesis. What changed in the team's tooling between sprint N and sprint N+1, and did the trend respond?
When those three sit together, you stop arguing about AI and start measuring it. The conversation moves from "are people using Cursor enough?" to "the team rolled it out three sprints ago, the health trend is up four points, and the per-developer AI spend rose twenty-two percent. We're up on net. Renew." Or — just as valuable — "the trend is flat, the spend is up, the maturity score didn't move. Try a different tool."
That's the closed loop: cost paired with outcome, on the same cadence. Not a dashboard with more numbers on it. A decision-grade picture that lets you stop guessing.
Per team, not per developer
A note that matters: this only works if the spend view is per-team, not per-developer.
The moment you put individual AI usage on a leaderboard, three things happen, all bad:
- Developers route around the measurement — running models locally, sharing accounts, opting out of the tools you're paying for.
- The data gets gamed. More prompts ≠ more value, but it'll show up that way.
- You've quietly become the surveillance product you said you'd never build.
Team-level spend paired with team-level outcomes is the right unit. It answers the budget question without creating the surveillance one.
The Question to Bring to the CFO
Don't bring an invoice. Bring two trend lines: team delivery health and team AI spend, on the same axis, over the last six sprints. The conversation changes immediately.
The Honest Move
If you're an engineering leader, you have two ways to handle the next AI budget conversation.
The first is to keep doing what most teams are doing: pull the invoice, narrate the vibe, hope the room buys it. This works for one or two cycles. It stops working the moment a peer team shows up with a trend line.
The second is to instrument before you advocate. Get the spend visible and the outcomes visible — not after the renewal, but before. You'll find one of three things:
- The spend is producing durable outcomes. Defend it loudly. You now have the data to do it.
- The spend isn't producing outcomes. Cut it before someone else cuts it for you. Better to volunteer the trim than be the case study.
- The spend is producing outcomes for some teams and not others. That's the most common answer, and the most useful one. Adoption isn't homogenous; tooling fit isn't either. Spend should follow fit.
In every case, you're acting on evidence. That alone separates you from most of the conversation happening in your industry right now.
The Bottom Line
In 2026, "is AI working?" stopped being a thought-leadership question and started being a budget question. The teams that win are the ones who can answer it on demand, with two trend lines, in under a minute.
Did it ship? Did it stick? And what did it cost?
The first two we wrote about in March. The third one is the half that finance is asking about now. The teams that have the answer aren't the ones using the most AI. They're the ones who can prove what they got for it.
Show me the spend. Show me the trend. Now we can talk.
Measure developer effectiveness, not just productivity
Six dimensions of effectiveness. Trends over time. Insights that help your team see what's working.
Further Reading
- Ship and Stick: Measuring Whether AI Is Actually Working — the outcome half of the question, with the data behind why activity metrics are vanity.
- The Seven Deadly Sins of Engineering Metrics — anti-patterns to avoid when you instrument the loop.
- DORA Metrics Without the Dashboard Tax — why the data you need is already in the integrations you've already connected.
Continue reading
- Ship and Stick: How to Measure Whether AI Is Actually WorkingEvery organization is adopting AI. Almost none can prove it's working. Here's how to measure what actually matters — outcomes that ship and stick, not speed that breaks everything. · 12 min read
- Twelve Working Agreements for Machine-Written CodeThe vibe-coding hangover retro ends with rules on a whiteboard. Here are twelve you can steal — each one a single-sentence rule, the number that moves if it's holding, and the check-in that keeps it honest. · 13 min read
- The Retro for the Vibe-Coding HangoverAI made your team faster in week one and slower by month three. Code churn is up 861%, incidents are up 242%, and the fix isn't less AI. It's the ceremony you already run — fed with real data instead of vibes. · 9 min read