Book a demo

Cost per AI Outcome: How to Prove AI Value to Your CFO with the 5 Hows

FinOps

11 Mins Read

|

Michael Stephenson
Michael Stephenson Azure FinOps Practitioner

Quick answer: Cost per AI outcome is the total cost of an AI solution divided by the measurable business outcomes it delivers. The hard part is not the division. It is finding an outcome that ties to a business KPI and proving the AI caused it. The 5 Hows framework gets you from a business goal to a measurable outcome, and a control group proves the AI made the difference.

When I was at the FinOps X / Tokenomics event in Amsterdam recently one of the things that stood out was that many companies struggle to communicate the value of their AI solutions.

The general sentiment was:

“How do we prove ROI on AI in terms our CFO will actually accept?”

At the event we were talking about “Cost per Outcome”.

This is great but I think at the current maturity level in AI Economics we are struggling to get from this abstract concept to something specific.

Download the free Claude Skill and use the 5 Hows framework on your next AI initiative

Get the Claude Skill

Why is AI ROI so hard to prove to a CFO?

For your AI solution the problem you have is:

  1. What is your outcome
  2. Is the outcome measurable
  3. Does the outcome actually tie back to a KPI
  4. Is X number of the outcome translatable to something the CFO would actually care about

What is cost per AI outcome?

Cost per AI outcome is the total cost of an AI solution divided by the number of measurable business outcomes it delivers over the same period.

Cost per AI outcome = total AI solution cost ÷ measurable outcomes attributable to the AI

The formula itself is simple. The challenge is working out what the outcome actually is and proving the AI delivered it.

I think this is a similar problem to articulating the value of your cloud solutions and in the FinOps world we have the Cloud Unit Economics concept to help us describe the value in a measurable way.

I think an AI solution is really just a variation of this problem and that Unit Economics should be able to help us describe the cost per unit in a meaningful way to the CFO.

The thing that I think AI maybe makes more complex is “How do we get from an outcome to a unit?”

How do you find a measurable AI outcome? The 5 Hows framework

While I was listening to some of the talks covering this challenge of identifying the value of my AI solution, for some reason the 5 whys technique popped into my head. The 5 whys technique is normally used to help you identify the root cause of an issue but I was wondering if I could adapt this and use something like asking the question “How?” 5 times to help me identify the value indicator I am looking for.

Start with one Why

Before jumping into 5 hows, let’s start off by asking 1 why.

“Why are we building this agent or AI solution?”

Normally for your organisation this would map to one of a few macro level reasons such as:

  • Revenue Growth
  • Cost Reduction
  • Risk Reduction
  • Environmental Sustainability
  • Employee Experience

If I ask why are we building this and get 1 or 2 top level reasons then I can ask how. Below I’ll look at an example.

Worked example: a churn reduction agent

The churn reduction agent is an AI agent that will be built to help our customer success team reduce churn. It is going to cost $10,000 per month to run. How do we measure the value of this agent?

To start with at the macro level the Churn Reduction Agent will improve our Revenue Growth. This is the reason we are building it.

  1. How will the agent improve your revenue growth?
    1. It will help reduce the number of customers who churn
  2. How will the agent help reduce the number of customers who churn?
    1. It will identify at-risk customers earlier and intervene before they leave
  3. How does it identify “at-risk”?
    1. It monitors behavioural and usage signals such as declining logins, negative sentiment, unresolved tickets and scores them against a churn-risk model
  4. How does it intervene?
    1. It triggers an action to alert a Customer Success Manager (CSM) with context and a recommended personalised playbook
  5. How will we know the agent intervention changed the outcome?
    1. We will compare the retention rate for flagged accounts that received an intervention vs a control group where the account was flagged as at risk but an intervention was not triggered.
  6. How will we then convert this change of outcome to a measurable metric?
    1. We can look at the number of accounts saved in the control group vs the group where the agent intervened and then multiply this by the average annual revenue for these customers.

At this point you might have noticed I asked 6 Hows. In this case I needed one more to get to something measurable. I think asking exactly 5 is not the point, it’s more that it sends you down a path to get to the value.

I think in the real world your AI solution might target more than one macro level metric and you might ask multiple hows for each one.

The point isn’t to ask exactly 5, it’s to keep digging in a structured way until you get to the specific measurable outcome.

We have now transitioned from “Cost per outcome” to “Cost per measurable outcome”.

How do you prove the AI caused the outcome?

At this point we know how we will identify that the agent added some value and how to compare it against scenarios where the agent did not get involved. We now need to express this as a defensible unit.

We could take the following metrics:

  • 200 accounts intervened by agent
  • 30 of those still churned → retained = 170 → retention rate (intervened) = 170/200 = 85%
  • 200 accounts identified but not intervened (control group)
  • 70 of those churned → retained = 130 → retention rate (not intervened) = 130/200 = 65%
  • Average ARR per account = $10,000

Retention Lift

65% 85% = 20 percentage points

The agent improved retention by 20 points over doing nothing.

Incremental Accounts Saved

20 percentage points × 200 intervened accounts = 40 accounts

Of the 170 retained accounts, only 40 are actually attributable to the agent — the other 130 likely would have stayed anyway, based on the 65% baseline.

How do you calculate cost per AI outcome?

Monetising the Value

40 accounts × $10,000 ARR = $400,000 incremental ARR saved in the agent group.

Divide cost by outcomes

This gives us our cost per measurable outcome: year 1 cost of $170,000 ($50,000 build + $120,000 running) ÷ 40 saved accounts = $4,250 per retained account, against $10,000 of ARR retained for each one.

How do you capture the full cost of an AI agent in Azure?

Before you can divide cost by outcomes, you need to know what the AI solution actually costs. In Azure, the cost of an agent is rarely a single line on the bill. It is usually spread across Azure OpenAI or Azure AI Foundry model usage, the compute and storage the agent runs on, and the data services that feed it signals, such as the usage and sentiment data our churn agent depends on.

If these resources sit in shared subscriptions or resource groups, the agent’s true cost gets lost. Tagging the resources that belong to the agent and allocating shared costs fairly gives you a cost figure you can defend. Turbo360 helps here by allocating Azure costs to a specific workload or business unit and tracking how that cost changes month to month, so your cost per outcome stays accurate as usage grows.

The cost side should include the time Customer Success Managers spend acting on the agent’s alerts, not just the agent’s running cost. If each intervention takes an hour of CSM time, that is part of the true cost per measurable outcome.

What’s a good cost per AI outcome?

There is no universal benchmark for cost per AI outcome, because outcomes are worth very different amounts to the business. A retained customer, a resolved support ticket and a qualified lead all carry a different value.

A better test is to compare the cost of each outcome with the value of each outcome. In the churn example, each saved account costs $4,250 and retains $10,000 of ARR, around a 2.35x return in year 1. The agent breaks even after 17 saved accounts and pays back in around 5 months.

A higher cost per outcome can be perfectly acceptable if each outcome is valuable. A low cost per outcome means very little if the outcome doesn’t tie back to a KPI.

How do you present AI ROI to your CFO?

At this point we can now say to the CFO:

We invested $50k to build an AI solution which will have an annual running cost of $120,000 per year. Based on a 12-month pilot comparing 200 at-risk accounts the agent acted on against a 200-account control group, the outcomes we expect are:

  • Our retention of at-risk customers will improve from 65% to 85%
  • Churn among at-risk customers will fall from 35% to 15%, a 20 percentage point reduction
  • The agent needs to save 17 accounts in year 1 to break even for running cost and build cost. Assuming the churn is evenly split across the year the payback period would be 5 months
  • On a group of 200 customers the agent acted on there was $400,000 in ARR that would have been lost compared to the group where the agent didn’t act.
  • Each account the agent saves costs us around $4,250 against $10,000 of ARR retained.

Download the free Claude Skill and use the 5 Hows framework on your next AI initiative

Get the Claude Skill

How do you measure AI outcomes over time?

  1. In year 2 the number of at-risk customers will potentially go down because you would hope that the agent has helped you to reduce the number of customers who are at risk on their next renewal. This may be an additional metric you want to bring in longer term as justifying the year 2 and beyond value may be calculated slightly differently.
  2. If the agent pilot is approved for full roll out then the larger customer base is likely to bring in more saved revenue as the agent was only acting on a smaller subset of customers at risk of churn
  3. As time progresses there may be other factors that come into play which affect the churn comparison. In the real world you are normally in a situation with multiple variables to consider and over a longer period the agent won’t be the only thing that affects churn. Maybe a new feature comes out which improves things or makes things worse. I think it’s good to do the side by side comparison with and without the agent to begin with but if the agent is successful then you will want to use it on your whole customer base and the result will be that you won’t have a control group where the agent is not involved to compare with so longer term evaluation of value may be affected by many things.

What mistakes distort cost per AI outcome?

  1. You would need to consider how the agent contributed to retention and the techniques it used. As an example if the agent just offered customers a big discount to keep the customer then this may affect the numbers. You might consider the ARR up for renewal before the agent intervened and the ARR retained after renewal so any discounts were factored in.
  2. A control group of 200 accounts is fairly small, so it is worth checking the retention lift is statistically meaningful before presenting it. There is also a business trade-off in deliberately not intervening with at-risk customers, so keep the control group as small and short-lived as you can while still getting a credible comparison.

Summary

In summary, I hope this article gives people some ideas for the initial stages of an AI project, where you need to work out the value your AI solution will deliver and express it in a measurable, CFO-defensible way. The approach is simple: ask one Why to anchor the solution to a business goal, keep asking How until you reach an outcome you can measure, compare against a control group to isolate the agent’s impact, turn that impact into money, and then divide your cost by it to get your cost per measurable outcome.

FAQ

How is cost per AI outcome different from cost per token?

Cost per token measures what the model costs to run. Cost per AI outcome measures what each business result costs. Token cost is useful for engineering teams, but outcome cost is what a CFO needs to decide whether the AI solution is worth funding.

Do you need to ask exactly five Hows?

No. The number isn’t the point. Keep asking How until you reach an outcome you can measure, compare against a baseline and convert into money. In the churn example, it took six.

How big should a control group be?

It depends on the size of the improvement you expect. A large lift can be proven with a smaller group, while a small lift needs more accounts. Balance statistical confidence against the business cost of not intervening with at-risk customers.

Can you measure cost per AI outcome without a control group?

Yes, you can use a historical baseline instead, such as churn rates before the agent launched. It is weaker evidence, because other factors change over time, but it is a practical option once the agent is rolled out to your whole customer base.

Advanced Cloud Management Platform - Request Demo CTA

Related Articles