TL;DR: Twenty widely shared AI ROI metrics measure revenue, cost, cycle time, and adoption. Not one measures risk. Nothing counts incidents avoided, blast radius contained, or agent actions blocked. A business case with no risk column looks great until the first incident, and then it stops looking like a business case at all.
Clare Kitching published an infographic on LinkedIn called "20 ways to measure AI impact." It's a good piece of work. Twenty metrics, each with a formula, grouped into six categories: financial, operational, customer, scaling, workforce, and adoption.
Read all twenty and something is missing.
What do the standard AI ROI metrics actually measure?
They measure benefit and they measure uptake. The financial five cover revenue uplift, net cost savings, payback period, benefit-cost ratio, and cost per AI-assisted transaction. The operational three cover cycle time reduction, straight-through processing value, and labour capacity released. Customer metrics track conversion, retention, satisfaction, first-contact resolution, and collection efficiency. Scaling metrics track time to value, the share of pilots that reach production with proven value, and how many priority value pools AI touches. Workforce is employee experience. Adoption is active users, business units engaged, and an AI literacy score.
That's a serious framework. It's the shape of an AI business case as most companies build one today.
Now count the metrics that measure what happens when the AI does something wrong.
Zero.
What's missing from every one of these frameworks?
The risk column. There's no metric for incidents avoided, no metric for blast radius contained, no metric for audit findings closed, and no metric for agent actions blocked before they executed.
That absence isn't specific to one infographic. It's how the whole category is currently framed. AI investment gets measured the way a productivity tool gets measured, because that's the frame the technology arrived in. An AI agent isn't a productivity tool. It holds credentials, calls systems, and takes actions at machine speed on your behalf.
This shows up in the market too. Cost and ROI briefs are now a confirmed recurring genre in my own call ledger, three in a five-week window. Buyers are actively trying to build these cases. They're building them without the column that would tell them what the downside costs.
Why does a missing risk column break the business case?
Because it makes the return look one-directional.
Every metric on the list moves in one direction. Revenue goes up, cost goes down, cycle time shrinks. There's no term that can go negative. So the model can't produce a bad outcome, which means it can't be wrong, which means nobody stress-tests it.
Then an agent does something nobody scoped. Now there's an incident cost, a remediation cost, an audit finding, and in a regulated environment a reporting obligation. None of those were in the model. They don't reduce the measured ROI, because the model has no line to put them on. The number on the slide stays green while the actual return goes negative.
The uncomfortable version: a company with no governance and a company with mature governance will report the same ROI on this framework, right up until one of them has an incident.
What about the baseline problem?
Every metric on the list is a difference. Post-AI revenue minus baseline revenue. Baseline cycle time minus post-AI cycle time. Post-AI CSAT minus baseline CSAT.
Each one assumes you have a clean baseline. Almost nobody does.
Most companies deployed AI into processes they had never instrumented. There's no clean pre-AI measurement of cycle time on the process the agent now runs, so the baseline gets reconstructed after the fact from memory and rough estimates. A difference between a measured number and a remembered number isn't a measurement.
If you're going to build the case, build the baseline first, and write down how you measured it. That single act does more for the credibility of the number than any of the twenty formulas.
Who verifies a "verified" benefit?
One of the twenty metrics is the AI benefit-cost ratio: verified AI benefits divided by total AI costs.
That word "verified" carries the whole metric. Verified by whom, against what evidence, on what cadence?
In practice it usually means the team that built the thing reported that it worked. That's attestation, not verification. The same question sits underneath agent governance generally: an agent that reports its own success is not evidence, and a log that the agent can write to is not an audit trail.
If a number in your AI business case is labeled verified, name the verifier in the same sentence. If you can't, call it reported and move on.
What should you add to the model this week?
Five lines. None of them need a new tool.
Incidents avoided. Count the agent actions your controls blocked. If the answer is zero because nothing blocks anything, that's the finding.
Blast radius per agent. For your most active agent, list the systems it can reach and the actions it can take. Cost the worst plausible chain, not the worst single action.
Time to contain. How long from an agent misbehaving to it being stopped. If nobody can answer in minutes, the number is hours, and hours at machine speed is a lot of actions.
Governance drag on pilots. Two of the twenty metrics track pilot-to-production rate and value coverage. Governance is usually what holds a pilot at the gate. Measure how long the gate takes, because that's a real cost of not having a standard.
Cost of the ungoverned baseline. What you'd spend to reconstruct evidence for an auditor about what your agents did last quarter. If that number is large, it's a liability sitting off the model.
Put those next to the twenty and the case gets harder to write, and much harder to knock over.
Where does Zero Trust fit here?
Zero Trust is the foundation this sits on. It verifies every request against identity, device, posture, and behavior signals, continuously, per request rather than per session. That's what makes any of the risk metrics above measurable in the first place. You can't count blocked actions without a control point that does the blocking.
What Zero Trust needs on top for agents is a view of the sequence. A single agent request can look completely normal while the run of fifty requests around it adds up to something nobody approved. The risk column only tells the truth if something is watching the trajectory, not just the transaction.
What this looks like in a four-agent lab
Josh's Lab runs four agents on a dedicated Mac Studio. Atti orchestrates, Forge codes, Scout researches, Quill writes.
The cost side of that lab was easy to measure from day one, because the token bill arrives whether you asked for it or not. One overnight run burned $300 when agents spawned more agents faster than anything counted them.
The risk side took real work to measure, and it needed a control point before it could be counted at all. Until something could refuse an action, "actions blocked" was structurally zero, and zero looked like safety instead of blindness.
That's the trap in miniature. A metric that reads zero because nothing is watching is indistinguishable, on a slide, from a metric that reads zero because nothing went wrong.
Key takeaways:
Twenty common AI ROI metrics cover financial, operational, customer, scaling, workforce, and adoption. None measure risk, governance, or security.
A model where every term moves in one direction can't produce a bad outcome, so nobody stress-tests it.
Every metric is a difference against a baseline, and most companies never instrumented the process before AI touched it.
"Verified benefits" is doing heavy lifting in these frameworks. Name the verifier or call the number reported.
Add five lines: incidents avoided, blast radius per agent, time to contain, governance drag on pilots, and the cost of reconstructing evidence.
Frequently asked questions
What is risk-adjusted AI ROI?
It's an AI business case that carries downside terms alongside benefit terms. Standard frameworks measure revenue uplift, cost savings, and adoption, all of which move in one direction. A risk-adjusted model adds incidents avoided, blast radius, time to contain, and the cost of reconstructing evidence, so the number can move both ways.
Why don't standard AI ROI frameworks measure risk?
Because AI arrived framed as a productivity tool, and productivity tools get measured on output. An agent that holds credentials and takes autonomous actions carries a different risk profile than a tool that drafts text, but the measurement frame hasn't caught up.
How do I measure incidents my AI agents avoided?
Count actions your controls refused. That requires a control point in the path that can say no and log the refusal. Without one, the count is structurally zero, which reads as safety on a slide and is actually an absence of visibility.
What is blast radius for an AI agent?
The set of systems it can reach and actions it can take if something goes wrong. Cost it as a chain rather than a single action, because agents work in sequences, and the sequence is usually worse than any one step in it.
Do I need new tools to add a risk column?
No. Start with a spreadsheet: your most active agent, the systems it touches, what it can change, and how fast someone could stop it. Most of the value is in discovering which of those four you can't answer.
What's the most common mistake in an AI ROI model?
A reconstructed baseline. Every metric is a difference, and most companies never measured the process before AI touched it, so the "before" number is a memory. Build the baseline first and write down how you measured it.
Twenty metrics tell you what AI gave you. None of them tell you what it can take away. Until the model carries a term that can go negative, it isn't a business case, it's a forecast of the good half.
Metrics framework credit: Clare Kitching, "20 ways to measure AI impact," LinkedIn, 2026.
