It found 5 per cent who say their data is adequately ready to support them. Both numbers come from the same companies, which means roughly nine in ten organisations are currently doing the thing they say their data is not ready to support. Sitting in the gap between those numbers is the most seemingly-respectable sentence in enterprise AI: we need to clean our data first. It sounds prudent. It comes with a project plan, a steering committee and a budget line. Nobody was ever sacked for proposing it. The readiness everyone is waiting for does not exist, and the admission is sitting inside the readiness literature itself. Data readiness is not a precondition you complete. It is a practice you start, and it starts the day a workflow goes live.

The politest way to say no

The readiness literature has an authorship problem. The data platforms publish readiness indexes because readiness is the product. The consultancies publish readiness frameworks because remediation bills by the month. The file-storage vendor discovered that nearly every enterprise is struggling to manage its unstructured data, and happens to sell the fix. The surveys can all be true and still be sales collateral with a methodology section. The discount applies to the survey up top, too; Dun & Bradstreet sells data. What survives the discount is the gap between its own two numbers, and the gap is the confession.

There are more confessions where that came from. When Cloudera and Harvard Business Review's research arm asked what actually blocks AI data preparation, siloed sources came first at 56 per cent and the absence of any clear data strategy came second at 44. Data quality ran third. And Gartner supplied the readiness decks' favourite stat: 60 per cent of AI projects lacking AI-ready data abandoned through 2026. That forecast was published in February 2025, on 2024 survey data, and its window shuts in December. It is still quoted as though the result were already in. The same firm's research says data is only ever AI-ready relative to a specific use case, then says the quiet part outright: there is no way to make data AI-ready in general or in advance. The firm that armed the stall keeps disarming it.

So when a remediation program asks for another two quarters, the question for the executive is what, specifically, the cleaning is for. If nobody in the room can answer, the cleaning has no customer, and the program is a deferral earning compound interest.

Ask the last question last

Watch how the question arrives in boardrooms, because everything downstream turns on the order. One board hands down a target: a dozen AI initiatives in the annual plan, progress reported by year end. Another ties AI to the efficiency program, so the use cases go hunting for costs, and the biggest cost in any knowledge business is people. Efficiency-framed mandates find efficiency-shaped answers. The mandate never says headcount, and the whole building hears it anyway.

Run the sequence in its actual order and it comes out differently. Why does the organisation exist, and where is the growth it has not been able to reach. What would it take to reach it, counting AI as one more tool in the kit alongside people, capital and time. And only then, scoped by that single opportunity: what data does this need. Asked in that order, the data question stops being an ocean and becomes a list. Sometimes the honest answer is that a workflow should not be automated at all, because it would spend trust the organisation cannot buy back. That answer is strategy too.

The data question is real. It is simply the last question, and the market has been asking it first. Where a regulator is watching, the last question also has teeth: the obligations that attach to a material process bind before go-live, and they attach to named workflows, never to an ocean nobody can certify.

The mess is the material

The technology changed sides. Every previous generation of enterprise software needed clean, structured input, because tables were all it could read. Large language models are the first enterprise tools that can actually read the other 80 to 90 per cent, the contracts, emails, meeting notes and policy drafts where the organisation lives. The warehouse keeps its job for the numbers; this argument is about everything it never read.

Three versions of the same policy used to be a data-quality defect. Now it is a finding: which one is current, who decides, and why do two departments believe different answers. The model surfaces the signal, a human adjudicates it, and the record improves while the work runs. That pairing only holds if the workflow shows its sources. The failed pilots are real, and most died the same way: a model pointed at everything, answering from anything with confidence, nobody assigned to adjudicate.

In our work at New Dialogue, the messy, governed start is the design: two or three sources, chosen and permissioned, one workflow, a human in the loop. Messy does not mean open. Autonomy comes later, after the record has earned it.

One more thing

Nirvana recorded their first album in thirty hours for $606.17 and named it Bleach, after a public-health poster Kurt Cobain had seen urging heroin users to clean their needles. The polish came later, and so did everything else. It took Aberdeen, the logging town Cobain spent his youth trying to leave, until 2005 to recognise him officially. The town did not use his name. It put four of his words on the welcome sign at the city limits, where they still stand: Come as you are.

Author: Matt Vitale