The CFO test for AI investments.
AI budgets are getting their first serious CFO scrutiny. The honeymoon period in which a slide on generative AI was enough to unlock funding is closing, and finance leaders are starting to ask sharper questions. The answers, in most boardrooms we sit in, are still too soft. This piece proposes a simple CFO test that any AI investment should pass before it gets renewed.
Why CFOs are tightening on AI specifically.
In the conversations we have with finance leaders, the pattern is consistent. The CFO function has been broadly supportive of AI investment in the first wave, often funding initiatives outside the normal business case discipline, on the understanding that the technology was new and the learning curve had value in itself. That tolerance is fading.
Several forces are converging. The headline costs of model inference are no longer the main driver, but the cumulative bill across multiple use cases and increasing usage is becoming visible. The first cohort of investments is approaching renewal, and finance teams want to know what was actually delivered against the original case. Boards are pushing back on the line item, sometimes because they themselves are being asked the question by investors. The narrative carry that AI enjoyed in the first cycle is being replaced by a normal scrutiny that any meaningful investment receives at maturity.
This is not bad news for the AI agenda. It is, in our experience, the moment at which AI investment moves from being an experiment to being a managed line of the operating model. The teams that prepare for that conversation early tend to come out of it with more credibility and, often, more budget than they expected. The teams that arrive unprepared lose ground.
The five questions.
The CFO test we use with clients and with our internal investment committees is built around five questions. They are simple to state and demanding to answer well.
1. What does this investment replace.
The first question we recommend asking, and the one most often dodged in early business cases, is what existing cost or work the AI investment displaces. If the honest answer is "nothing, this is purely additive", the case has to be built on revenue or option value, and the framing of the rest of the test changes. If the answer is "a category of work currently done by a named team or vendor", the conversation becomes much sharper, and the size of the prize becomes measurable.
2. What does it accelerate.
Some AI investments do not replace work but compress it. A cycle that took weeks now takes days. A decision that required several meetings now requires one. The value here is real but harder to instrument, and it belongs in the case as a named compression of a specific process, not as a vague productivity uplift across the function.
3. What is the unit economics post-deployment.
Once the system is in real use, how much does one inference, one workflow, one decision actually cost. Inference at scale, monitoring infrastructure, retrieval indexes, prompt maintenance and the engineering time to keep the system aligned with the underlying models all belong in this number. In the briefs we audit, this question is the one where the original business case usually breaks. Costs that were marginal in the pilot turn out to be the dominant line by the second quarter of real use.
4. What is the marginal cost of one more user or call.
Investments that look healthy in aggregate sometimes hide an unfavourable marginal cost curve. The next user, the next team, the next workflow does not necessarily cost the same as the average one. If marginal cost rises faster than marginal value, the case looks fine until the volume hits a threshold and the line item explodes. Asking this question explicitly is how a finance partner spots the cliff before the team walks off it.
5. What kills it.
Every investment has a set of conditions under which it stops making sense. A model deprecation, a price change from the underlying provider, a regulatory shift, an unexpected drop in usage, a competitor making the capability free. Naming these conditions upfront is not pessimism, it is hygiene. It tells the renewal committee, twelve months later, what to watch and when to pull the plug if the conditions trigger. The investments we have seen survive longest are usually the ones whose owner could answer this question crisply at the original approval.
The investments that survive renewal are not the ones with the smartest pitch. They are the ones whose owner could already answer the five questions on the day they asked for funding.
How to instrument the answers.
Asking the five questions is not difficult. Answering them with evidence rather than assertion is where the real work sits. From the engagements where we have helped structure this kind of business case, three practical disciplines tend to make the difference.
Treat the pilot as the instrumentation phase.
The pilot should not just prove that the technology works. It should produce the cost telemetry, the usage curves and the displacement evidence required to defend the case at renewal. That means deciding, before the pilot starts, what data will be captured and which lines of the P&L will be examined. In practice this is a conversation between the operational owner and the finance partner during the pilot scoping, not after it.
Distinguish what scales linearly from what does not.
In any AI investment, some costs scale per use (inference, retrieval, evaluation calls) and others scale with the surface area of the system (engineering maintenance, monitoring, governance). A business case that treats both categories the same usually misses the second one. A clean separation makes the conversation about future volumes much sharper, and the breakpoint conversation possible.
Bring finance into the case authoring, not the case reviewing.
When the finance partner first sees the case at approval, the conversation tends to be adversarial. When the finance partner has helped frame the case during pilot scoping, the same conversation becomes joint problem-solving. The numbers we see hold up at renewal are usually the ones that were stress-tested early, not the ones that were polished late.
When AI investment is genuinely defensive.
There is a category of AI investment where the five questions look harsh because the rationale is not productivity or revenue but defensive readiness. Building internal literacy. Avoiding being left behind by a competitor. Keeping the option of moving faster when a clearer use case emerges. These investments exist and they can be legitimate. The discipline is to argue them on their actual basis rather than dress them up as productivity cases.
In practice, the question we recommend asking is whether the investment is being justified primarily by cash flow effect or primarily by option value. If the answer is option value, the case should be funded under a separate envelope, with a different evaluation horizon and explicit criteria for when the option will be exercised. Mixing the two erodes finance trust faster than any other pattern we have observed, because at renewal the investment is judged on productivity criteria it was never structured to deliver.
A CFO who understands the difference between cash-flow cases and option-value cases is usually willing to fund both. A CFO who has been burned by option-value cases dressed as cash-flow ones tightens on everything, and the entire AI portfolio pays the price for the mislabelled ones.
Discipline of measurement as credibility asset.
The teams whose AI portfolios are growing fastest, in the boardrooms we observe, are not necessarily the ones with the most ambitious initial pitches. They are the ones whose finance partners have come to trust their numbers. That trust is earned over several renewal cycles by being honest about what worked, candid about what did not, and surgical about killing the cases that do not deserve another year of budget.
If your AI portfolio is approaching its first serious renewal cycle, the conversation we tend to have with clients in this position is less about choosing the right model and more about preparing the right case. The questions a sharp CFO will ask are predictable. Answering them well is a craft that compounds.
If you would like to compare notes on a portfolio review or on the case structure for an upcoming investment, the Consulting team can sit on that file directly. The brief form below opens the conversation and we respond within one working day.
Frequently asked questions.
What payback window should a CFO expect on an AI investment?
The payback window depends on the type of investment. For productivity-led use cases on internal teams, the expectation should sit within the same budget cycle. For revenue-led or customer-experience use cases, a longer window is reasonable, but the early indicators should be visible within the first quarter of real use. Investments that cannot articulate any payback window should not be renewed without a sharper case.
What are the hidden costs of AI that CFOs miss most often?
Inference costs at scale, evaluation and monitoring infrastructure, the engineering time spent maintaining prompts and retrieval indexes, and the human cost of operators learning new ways of working. None of these appear on the original business case. They become visible in the second or third quarter of real usage, by which point the budget has been committed.
How should opportunity cost be framed in an AI business case?
Opportunity cost belongs in the business case as a named line. What other investment would this team or this budget have funded if this AI initiative were not green-lit, and what does the comparison look like. Treating AI as exempt from opportunity-cost reasoning is one of the patterns that erodes finance trust the fastest.
Is there a place for option value in AI investment decisions?
Yes, but it must be argued explicitly. An AI investment that earns its place primarily through option value (capability building, future positioning, defensive readiness) should be funded under a different line than one that earns its place through immediate cash flow. Mixing the two leads to investments that are evaluated on the wrong criteria.
How do we align CFO scrutiny with operational AI teams?
By making the finance partner a co-author of the business case, rather than a final reviewer. The operational team brings the cost structure and the realistic use case. The finance partner brings the discipline of comparison and the language of P&L impact. Investments built this way pass scrutiny faster and survive longer.
What if the answers to the test questions are uncertain at the start?
Uncertainty is acceptable. Silence is not. The right discipline is to name the uncertain answers explicitly, to commit to instrumenting them during the pilot phase, and to schedule a checkpoint at which they will be revisited with real data. An investment can be approved on uncertain answers, but it should not be renewed on them.
Where this lands
How we'd take this further with you.
Consulting pillar
AI-Augmented Enterprise
From maturity diagnosis to use case prioritisation to durable adoption across the organisation.
Consulting pillar
Performance & Value
Measurement systems, KPI design and value tracking that actually move the business.
Consulting pillar
Strategy & Innovation Governance
Setting direction, sequencing portfolios, governing decisions that compound over time.
Writing is one thing. Shipping is the other. Selected work from the partners writing here.
See the work
