You Don’t Need All Thirty Years

I’ve spent a good part of my career arguing that organisations underestimate the value of the information they’ve accumulated over decades. I still believe that. But there’s a mistake that follows naturally from the argument, and I see it often enough that it’s worth warning against.

Once leaders accept that their historical information matters to AI, the instinct is to make all of it AI-ready. Every archive, every repository, every box in storage, back to the founding of the business. It sounds thorough. In practice, it is one of the most reliable ways to make sure nothing useful happens for years.

The boil-the-ocean trap

Estate-wide readiness programmes have a familiar arc. The scope is enormous, so the first phase is discovery. Discovery takes longer than planned, because nobody has a complete view of what the organisation holds. The business case is hard to write, because the benefits depend on AI use cases that haven’t been chosen yet. Meanwhile, the teams actually building AI tools get on with it using whatever data is easiest to reach.

The result is the worst of both worlds: a large, slow programme that hasn’t delivered anything, running alongside AI projects that are built on unprepared data anyway.

The problem isn’t ambition. It’s that “make our historical data AI-ready” isn’t a well-formed goal.

Readiness belongs to a use case

AI readiness is not a property that data has in the abstract. It’s a property that data has for a particular purpose.

Gartner puts this bluntly: There is no way to make data AI-ready in general or in advance, because readiness depends on how the data will be used[1]. The same records might be perfectly adequate for one application and dangerously inadequate for another. A document set that’s good enough to help staff find precedents may be nowhere near good enough to drive automated decisions.

That changes the question. Instead of asking how to prepare thirty years of information, ask which business questions are worth answering, and what information would be needed to answer them well.

Start with the question, then follow the records

Take a hypothetical insurer that wants AI to help claims handlers with complex cases. The useful question isn’t “how do we make the claims archive AI-ready?” It’s narrower. Which types of claim benefit most from historical precedent? Which records contain that precedent – claim files, policy wordings in force at the time, correspondence, settlement decisions? How far back does relevance genuinely extend? What condition are those records in?

The answer might be ten years of three document types for two lines of business. That’s a project with a scope, a cost and a measurable outcome. Thirty years of everything is not.

Once the question is clear, triage becomes much simpler. For each body of information, three tests do most of the work:

  • Value: Would this information materially improve the answer to a question the business actually cares about?
  • Viability: Can it be made reliable at reasonable cost, given its format, condition and how it was originally captured?
  • Permission: Are we allowed to use it for this purpose, given retention obligations, the reasons it was collected and its sensitivity?

Information that passes all three is where the effort should go first. Information that fails any one of them can wait, or may never need to be made AI-ready at all.

Don’t confuse recent with relevant

There’s a trap on the other side too. The obvious shortcut is to triage by date: prepare the last five years and ignore the rest. That feels pragmatic, but it can throw away exactly what makes historical information valuable.

Gartner’s definition of AI-ready data includes being representative of the use case – including the errors, outliers and unexpected events the AI needs to handle[1]. Rare events are rare. The last time a particular kind of dispute arose, a particular market shock hit a portfolio, or a particular clause was tested may have been fifteen or twenty years ago. If the record of that exists anywhere, it’s in the older part of the archive.

So the test is relevance to the question, not age. Sometimes the answer is recent material. Sometimes it’s one specific decade. Occasionally it really is the full history of a narrow document type.

The rest doesn’t disappear

Choosing not to make information AI-ready isn’t the same as ignoring it. Some estate-wide disciplines are still worth doing everywhere: knowing broadly what you hold and where, applying retention schedules consistently, and making sure that information you’re not preparing for AI isn’t quietly exposed to AI tools anyway.

What changes is the order of work. Readiness expands use case by use case, with each one funded by the value of the last. The organisation learns what good preparation looks like on a manageable scope before scaling it. And the parts of the archive that never earn their way into that process are identified for what they are: information to keep securely, or information to let go.

How Dajon approaches it

At Dajon, we’ve spent nearly thirty years looking after organisations’ records across their whole lifecycle: storing them, digitising them, managing their retention and securely destroying them. That gives us an unusually clear view of what sits in a typical historical estate, and of how little of it needs the same treatment.

We help organisations work out what they hold across physical and digital archives, and then match it against the questions they actually want AI to answer. The information that earns its place is captured and prepared to a standard the use case needs. The information that must be kept but won’t be used is stored securely and kept out of AI’s reach. And the information that has reached the end of its life is destroyed properly, with the evidence to show it.

The aim is not the largest possible readiness programme. It’s the smallest one that delivers real value, followed by the next.

I’ve long asked whether an organisation can really have an AI strategy without a historical data strategy. I’d add a clarification. A historical data strategy is not a plan to make every record ready for AI. It’s a way of deciding which ones are worth it.


References

  1. AI-Ready Data Essentials to Capture AI Value Gartner[↩][↩]