Last week I gave a keynote on enterprise data in the age of AI and the same question reached me from two directions in the same few days, once from a conference audience in Leeds and once from a finance team over a drink afterwards. Everyone’s investing in AI but hardly anyone could tell me what they’d actually got back outside of the really valuable personal administration and assisting in the meeting or note taking output.
There is a number I keep using because it is a striking result, 95% of enterprise generative-AI deployments have delivered zero measurable P&L impact (MIT NANDA initiative, 2025). Sitting next to it is a Gartner projection that 60% of AI projects unsupported by AI-ready data will be abandoned through 2026 (Gartner, 2025). Two different studies, same direction of travel.
The easy read is that the technology is overhyped. I don’t think that’s it. The tools mostly do what they say and also have the ability to achieve so much more in the future. The trouble is what we are pointing them at.
Adoption was never the hard part
KPMG’s 2026 Global AI in Finance report puts active AI use in finance at roughly 75%, up from about 30% two years ago. So adoption has stopped being the dividing line, almost everyone is doing something. What separates the firms getting a return is whether they can trust and evidence what the AI produces. The same report shows the “assurance-ready” firms reporting several times the error reduction of everyone else. Put plainly, the teams getting something back are the ones who could stand behind the output, not the ones who simply moved first.
That distinction matters more in finance than almost anywhere, because finance doesn’t get graded on insight alone. It gets graded on being right, and on being able to show why. An answer you can’t trace is worth very little when the auditor, the board or a regulator asks where it came from.
Which runs straight into a property of the tools themselves. It is well known and discussed that large language models aren’t fact engines, they generate plausible answers without any view as to what is true. Point that at clean, well-defined, owned data and it’s genuinely useful. Point it at the same tangle of inconsistent definitions and unowned data that already made month-end painful and it gives you confident, well-formatted answers that happen to be wrong. The confidence is the dangerous part.
The sequencing is backwards
Here is the bit I find genuinely odd. Deloitte’s Q4 2025 CFO Signals has 54% of CFOs prioritising AI-agent integration, slightly ahead of the 52% prioritising data-quality improvement. So the agents are being put to work on the very foundation they depend on, before that foundation has been checked. We’re automating on top of data we already know we don’t fully trust, and quietly hoping the automation sorts it out.
Unfortunately it won’t. A new layer of intelligence sits on the same foundations as everything else: the data flows, the master data, clear ownership, and the definitions people actually agree on. Get those right and AI has something solid to stand on. Skip them and you’ve scaled the mess, faster and with more conviction than before.
Whenever I make this point, someone asks a fair question back. Isn’t AI supposed to help clean the data? Yes it can support us in the process and ultimately help govern and improve over time, but it can’t decide what a cost centre is for in your organisation, or which of two systems owns the headcount number, or what “active customer” means across three teams that each define it differently. Those are judgement calls about how the business runs. No model makes them for you.
What I’d actually do first
I’m wary of turning this into a checklist, because the honest answer is that it depends on where your foundations are now although the overall shape is consistent.
Start by being specific about the decision or process you want the AI to improve. Not “adopt AI” but the actual outcome you’re after. Then name the handful of data domains that decision genuinely depends on, and look hard at them, the definitions, the ownership, whether the data moves between systems on its own or by hand. That’s usually where the readiness question gets answered, often uncomfortably. Only then is it worth wiring the clever layer on top.
This is the readiness question sitting underneath the diagnostic I’ve been writing about and it’s the one I think more finance teams should ask before they commit a budget rather than after. Not “are we doing AI” but “will our foundations hold the weight of it.” The maturity of those foundations tends to decide whether the investment compounds or quietly evaporates.
None of this is an argument against AI in finance. I think the upside will show a real and sizable return and the teams that get the foundations right first will pull away from the ones that didn’t. That’s rather the point. I read that 95% less as a verdict on the technology and more as a verdict on what we built it on.
If your own foundations are the thing you’re not quite sure about, that’s the conversation I find most useful to have.
Neil
The short version of the Finance Operating Impact Diagnostic gives finance leaders a fast, self-scored read on whether their foundations will hold the weight of what’s next. financediagnostic.netlify.app