Having Data Is Not the Same as Having AI-Ready Data

Eighty-eight per cent of data and analytics leaders say their organisation has the data readiness it needs for AI. Forty-three per cent of the same group name data readiness as the single biggest obstacle to aligning AI with their business objectives [1].

Same survey. Same 505 leaders. Same year.

That contradiction is not carelessness on the part of the respondents. It is what happens when a question gets answered at two different levels of resolution. Asked in the abstract, most leaders are thinking about volume: we hold decades of records, we have a data warehouse, we have a cloud tenancy, of course we have data. Asked in the specific, they are thinking about a particular project that has stalled because nobody can establish where a set of documents came from, what they mean, or who is allowed to feed them into a model.

Both answers are true. They are just answers to different questions. And the gap between them is where most enterprise AI programmes are currently sitting.

Possession and readiness are separate conditions

The pattern repeats across every serious study I have read in the past two years.

IBM’s Institute for Business Value surveyed 1,700 chief data officers and found that 81% now report their data strategy is integrated with their technology roadmap and infrastructure investments, up from 52% in 2023. Yet only 26% are confident their data can support new AI-enabled revenue streams, with accessibility, completeness, integrity and consistency named as the barriers [2]. Strategy has matured considerably. Confidence in the underlying material has not moved with it.

Gartner found much the same thing earlier: 63% of organisations either do not have, or are unsure whether they have, the right data management practices for AI, based on a survey of 1,203 data management leaders. Its analysts predicted in early 2025 that through 2026 organisations would abandon 60% of AI projects unsupported by AI-ready data [3]. Whatever the eventual figure, the diagnosis has held: traditional data management is too slow, too structured and too rigid for what AI systems demand, and most organisations lack the metadata even to assess whether their data qualifies.

None of this is a technology problem. It is a definitional one. We have been using a single word – data – to describe two conditions that behave very differently.

The UK government has now written the distinction down

In January 2026, the Department for Science, Innovation and Technology published guidance on preparing government datasets for AI. Buried in the self-assessment section is a sentence that should be pinned above every AI steering committee: “the availability or openness of a dataset does not equate to AI readiness”, and many datasets remain unfit for AI use without targeted remediation [4].

The guidance is candid about how the estate got that way. Data collection has historically prioritised operational delivery, reporting and compliance, while reusability for analytics or AI received less attention. The result is information that sits in departmental silos with inconsistent documentation and insufficient metadata, difficult to repurpose. It also cites the National Audit Office’s warning that handing over raw data without information on quality or provenance invites misunderstanding and misuse, particularly by systems that learn patterns at scale without any contextual awareness.

Read that as a description of the public sector if you like. I would read it as a description of almost every large information estate I have worked with in thirty years, in banking, insurance, law and property alike. Records were captured to run a process, satisfy a regulator or settle a dispute. Nobody captured them so that a language model could reason over them in 2026.

Six tests that separate stored from ready

The useful move is to stop arguing about the phrase and start applying tests to a specific collection. These are the ones that consistently expose the gap.

  • Can the system locate it? A Cloud Security Alliance study commissioned by Thales found 56% of organisations have only partial visibility into where their data is actually stored [5]. Nothing downstream matters until this is answered.
  • Can a machine read it? A scanned image of a contract is a picture of a contract. Without an extraction pipeline, it contributes nothing to a retrieval system beyond a filename.
  • Is the provenance known? Where the record came from, how it was produced, and what it was originally for. Provenance is what allows an output to be defended rather than merely produced.
  • Do rights and sensitivity travel with the content? This is the one that catches programmes late. Permissions attach to the source file; extraction, chunking and embedding routinely strip them away, and the restriction quietly stops applying at precisely the moment the content becomes easiest to surface.
  • Does the context survive? The National Archives makes the sharpest version of this point in the DSIT guidance: it deliberately preserves how records were organised by whoever created them, because meaning comes partly from the relationships between records. Flatten a filing structure into a document lake and you destroy information that was never written down in any single file.
  • Can you trace an answer back to its source? If you cannot get from a model’s output to the specific document that produced it, you have a plausible system rather than an auditable one.

Fail any one of these and you still have the data. You do not yet have AI-ready data.

Why the historical estate is the hardest part, and the largest prize

Two details from the government’s departmental interviews stayed with me. DEFRA still operates more than 500 paper form services. And The National Archives raised concerns about digitising historical records without metadata, on the grounds that it creates governance and risk gaps.

That second point deserves emphasis, because it cuts against the instinct of every organisation that has ever approved a scanning budget. Converting a paper archive to PDF does not make it AI-ready. It makes it digital. If nothing captures what each document is, when it was created, which matter or client or policy it belongs to, and who may see it, then digitisation has moved the problem from a basement to a server and made it faster to retrieve the wrong thing.

The guidance’s own mitigation is worth borrowing: use AI to create structured surrogates while keeping the originals immutable and the decisions reversible. Enrich, index and contextualise the historical estate rather than transforming it in place. The record stays the record. What you build alongside it is the layer that makes it usable.

This is also why the historical estate is the prize rather than the burden. Current systems tell an AI what the organisation is doing. Only the historical record tells it what the organisation has learned – which claims turned bad, which clauses were negotiated away, which decisions were taken and on what basis. That knowledge is not recoverable from live data.

The question to ask before the next platform decision

The DSIT guidance reframes the whole exercise, and I think the reframing is the most valuable thing in it. The test moves from do we have the data? to whether a given collection can responsibly sustain AI use across its full lifecycle.

Gartner makes a complementary point: AI-ready data is not a one-off achievement but a practice, requiring continuing investment in metadata management, observability and governance. Readiness is a verb.

If you are weighing an AI investment this quarter, the practical starting move is small. Take one intended use case and one specific body of historical information, and run the six tests above against it honestly. You will learn more from that exercise than from another platform evaluation, and it will tell you which of the two things you actually have.

Because you can hold thirty years of institutional knowledge and still be unable to use a single page of it. Which returns us to the question underneath all of this: can you really have an AI strategy without a historical data strategy?

How Dajon helps

We spend our time in the part of this problem that vendors tend to skip: the decades of records organisations already hold. That means digitising with metadata rather than without it, preserving the structure and provenance that make historical information intelligible, and building the classification and access layers that let AI systems use it safely. If you want to know whether a particular archive is ready, that is a question we can answer concretely.


References

  1. 2026 State of Data Integrity and AI Readiness Drexel LeBow / Precisely[]
  2. IBM Study: Chief Data Officers Redefine Strategies as AI Ambitions Outpace Readiness IBM Newsroom[]
  3. Lack of AI-Ready Data Puts AI Projects at Risk Gartner[]
  4. Guidelines and best practices for making government datasets ready for AI DSIT[]
  5. Unstructured Data Surges as Enterprises Struggle to Maintain Visibility and Security Cloud Security Alliance[]