The Number Is Either Right or It Is Not: Why AI Belongs in Finance First

The Number Is Either Right or It Is Not: Why AI Belongs in Finance First

The gap that has not closed

In August 2026, McKinsey published its annual global AI survey: 1,719 respondents across 97 countries, weighted by national GDP contribution. Eighty-nine percent reported using AI regularly in at least one function. Forty-four percent were scaling it across the enterprise, up from thirty-eight the year before.

And thirty-seven percent reported any contribution to EBIT.

That number was thirty-seven percent the year before as well. Adoption rose. Scaling rose. The earnings line did not move.

This is the finding that should organise a board conversation, and it is more useful than the figure that gets quoted instead. You have probably seen the claim that ninety-five percent of AI pilots fail. It comes from a July 2025 report by MIT NANDA, and it does not mean what it is usually taken to mean: the denominator was all surveyed organisations, most of which had never piloted a custom tool at all, and the sample was 52 interviews and 153 conference-recruited survey responses. Among companies that actually ran a pilot, roughly a quarter succeeded. Meanwhile Menlo Ventures found that AI purchases convert to production at forty-seven percent, against twenty-five percent for conventional enterprise software.

So AI tools do reach production. What fails is the translation from a working deployment into a number the CFO can find. That is not a technology failure. It is a measurement and process failure — and it is the reason we think the finance function is the right place to start, not the last place to get to.


Why finance is the honest test

Every other function offers somewhere to hide. A marketing team can point to content volume. A support team can point to deflection rate. An engineering team can point to lines of code. None of these is a lie, and none of them is an earnings impact.

Finance has no equivalent. A reconciliation either ties or it does not. A variance is either explained or it is not. A forecast is either within tolerance or it is outside it. When you deploy AI into a finance process, the question of whether it worked is answered by the close calendar, not by a survey.

That discipline is uncomfortable, which is precisely its value. Consider the strongest piece of evidence we have on self-reported productivity: a randomised controlled trial published by METR in July 2025. Sixteen experienced developers worked through 246 real issues in mature codebases, with and without AI assistance. They expected to be twenty-four percent faster. They were nineteen percent slower. And after the fact, having been slowed down, they still believed they had been sped up by twenty percent.

A thirty-nine point gap between perceived and actual performance. That is what happens when the measurement instrument is a person’s impression. Finance does not permit that instrument.


Where the money actually is, and where the budget actually goes

The MIT report is weak on its headline and strong on a detail almost nobody quotes. Roughly seventy percent of generative AI budgets went to sales and marketing. The returns the researchers could actually locate were in the back office: document processing, customer service handling, reconciliation-type work, external creative spend that stopped being necessary.

McKinsey’s 2026 function-level pattern points the same way. Cost reductions are reported most often in supply chain, service operations and manufacturing. Revenue gains are reported most often in marketing and sales — but reported revenue gains are the softest category in the whole survey, and the one where attribution is hardest to defend.

The uncomfortable reading is that organisations are spending where the story is easiest to tell and finding value where the story is hardest to tell. Finance sits in the second group.


Three ways the gain disappears before it reaches the P&L

BCG published an analysis in August 2026 of why pilots that work still produce nothing. Its three mechanisms match what we see in Gulf deployments.

Automation without redesign

A task that took ten days now takes one. The customer still waits ten days, because the nine days sat elsewhere in the process. The AI worked. The process did not change. This is the most common failure we encounter and it is entirely a management problem.

McKinsey quantifies the corrective: roughly seventy-three percent of the firms it classifies as high performers fundamentally redesigned workflows around AI. Among everyone else, twenty-five percent did.

Efficiency gets reabsorbed

A twenty percent speed-up on an analyst’s work becomes twenty percent more analysis, not twenty percent less cost. Whether that is good or bad depends on whether more analysis was what you needed. What it is not is a structural cost reduction, and it should not be presented to a board as one.

Activity is measured instead of money

Hours saved. Tickets deflected. Documents processed. Every one of these is a real quantity and none of them appears in a set of accounts. If the business case was built on activity metrics, the review will be too, and the earnings question will never actually be asked.


The design decision that makes finance AI defensible

There is a technical choice underneath all of this that determines whether a finance deployment survives its first audit.

A language model is a probabilistic system. Asked to produce a number, it will produce a plausible number. In most functions the cost of a plausible-but-wrong output is embarrassment. In finance it is a restatement.

The way through is to separate the two jobs. Detection and calculation are deterministic work: rules, thresholds, comparisons against a defined semantic layer where net sales means one thing across every system that touches it. Explanation and ranking are language work. The model reads what the deterministic layer produced, judges what matters most, and writes it in a sentence a human can act on.

Figures are passed in. They are never generated.

This is the architecture we built RAIL on, and the reason is not elegance. It is that every figure has to trace back to the row that produced it, or the finance team will not sign it — and they are right not to.


What to ask before you fund anything

  • What line in the accounts does this move, and who owns that line? If no one owns it, the project has no reviewer.
  • What is the number today? Cycle time, hours, error rate, cost per transaction. Measured, not estimated. If you cannot state the baseline, you cannot claim the improvement.
  • What changes in the process, not just the task? If the answer is nothing, expect the ten-days-to-one-day outcome above.
  • Where does the efficiency go? Into reduced cost, redeployed capacity, or higher volume. Decide before, not after.
  • Can every output be traced to source? In finance this is not a nice-to-have. It is the difference between a tool and a liability.
  • Who measures it, and when? A date in the calendar, not a review in due course.

None of these questions is about AI. That is the point. The organisations reporting real EBIT contribution are not the ones with better models — everyone has access to the same models. They are the ones that treated an AI deployment as an operating change with a financial owner, and measured it the way they would measure any other capital decision.


Sources: McKinsey, The State of AI: Global Survey (25 August 2026); BCG, Why AI Pilots Rarely Deliver Real Business Value (31 August 2026); METR, Measuring the Impact of Early-2025 AI on Experienced Open-Source Developer Productivity (10 July 2025); Menlo Ventures, The State of Generative AI in the Enterprise (9 December 2025); MIT NANDA, The GenAI Divide: State of AI in Business (July 2025).

Leave a Reply

Your email address will not be published. Required fields are marked *