The Plan Nobody Has Made Yet

You may not agree with this. But I am going to say it anyway

Historical data is your organisation’s institutional knowledge. If you want AI to make decisions based on what your organisation knows – rather than what a model happens to know – that historical data matters enormously. What you make available to AI determines what organisational context it can draw on, and that in turn shapes whether the answers it produces can genuinely be trusted or whether they simply sound convincing enough to make it into the boardroom.

I don’t think enough organisations have made that connection yet. I say that because I ask them.

It’s one of the first questions I put to any organisation that tells me they’re investing heavily in AI. Not which platform they’ve chosen, or which model, or what the implementation timeline looks like. I ask what they’re doing about their historical data.

The answer, more often than I’d like, is some version of “we don’t have any plans for that yet.”

I put the question to a CIO recently. No plan for the historical data. A seven-figure AI budget approved shortly before. That, right there, is the disconnect.

I’ve worked in this industry for nearly thirty years, and in that time I’ve watched organisations change the way they use information several times over. One assumption, though, has been remarkably difficult to shift: historical data is something you keep because you have to. That assumption made sense once. I don’t think it makes sense any more.

Thirty years, three eras, one assumption

In the early 2000s, the relationship organisations had with historical information was straightforward: keep it. Store the files, archive the emails, retain the contracts. The question that mattered was “how long do we have to keep this?” not what could we learn from it, or what decisions it could improve, or what thirty years of it might tell us about the organisation. The focus was retention. Historical data was a compliance obligation and a cost to be managed, and that was largely the end of the thinking.

Then the conversation changed. Digitisation arrived, electronic document management went mainstream, and organisations scanned documents, captured invoices and took paper out of their operational processes. The emphasis moved from store it to use it – but mainly for whatever process was happening today. The historical archive stayed where it was, treated as something separate: retained, protected, and eventually disposed of according to policy.

Now we’ve entered another era entirely, and the information organisations accumulated over the previous thirty years suddenly has a completely different potential value. The trouble is that many organisations are entering the new era with the data strategy of the old one.

AI doesn’t automatically know your organisation

There’s a distinction in the AI conversation that I think sometimes gets lost. Modern AI systems arrive with extraordinary amounts of general knowledge. What they don’t arrive with is any knowledge of you.

They don’t inherently know why your business made a particular decision fifteen years ago, or what happened afterwards. They don’t know the exceptions that changed a policy, the claims that went wrong, the contracts that produced unexpected consequences, or the precedents your longest-serving people remember because they lived through them. They don’t know the correspondence that explains not only what happened, but why. That knowledge exists inside the organisation – and much of it exists in historical records.

If those records are inaccessible, poorly classified, stripped of context or simply excluded from the systems AI can retrieve information from, then the AI is working with an incomplete view of the organisation. I’m not claiming every AI error comes down to missing historical data; these systems can produce incorrect or misleading output for plenty of reasons. But asking an AI to make organisation-specific judgements without giving it access to the relevant organisational knowledge creates a fairly obvious problem. It can only reason from what it can actually reach.

Which leads to a distinction I’d want every organisation investing in AI to sit with for a moment: having historical data is not the same as having AI-usable historical data.

Thirty years of data can still mean very little to AI

An organisation can hold decades of information and still have very little of its institutional knowledge genuinely available to AI. The information might be spread across old systems. Some of it may still be physical. Some was scanned years ago with barely searchable text and no meaningful metadata. Legacy records sit in proprietary formats, naming conventions drifted, classification schemes changed several times, and records that arrived through mergers and acquisitions follow structures nobody currently in the building designed. Perhaps most importantly, the context that gives those records their meaning may never have been captured in a structured way at all.

The information exists. Existence is not usability.

This is also why I don’t think the answer is simply “more data.” Volume alone doesn’t create intelligence. A million badly classified documents with no context are not necessarily worth more than ten thousand well-governed, relevant records. For historical data to become useful organisational knowledge, AI needs to be able to find the right information, understand its relationship to other information, and use it within the appropriate permissions and governance controls. That takes work – and in many organisations, that work isn’t currently part of the AI programme at all.

Then someone plays the compliance card

This is usually where the conversation gets interesting.

I ask what the plan is for the historical data. Someone says they’ll need to check with compliance. Compliance says it needs to go to legal. Legal is looking at it. The momentum quietly disappears – while the AI programme carries on regardless.

I do understand why. Historical data sits at the intersection of legal, compliance, technology, information governance and operations, and ownership is rarely clear. Different categories of information carry different retention requirements, privacy considerations, access controls and commercial sensitivities. These are real issues, and they shouldn’t be dismissed.

But compliant and ready are not the same thing. An organisation can comply fully with its retention obligations and still have a historical data estate that is extremely difficult for AI to use effectively. Retention answers what must be kept, why, and for how long. AI readiness asks different questions altogether. Can we find it? Can we extract anything useful from it? Do we understand what it represents, and how it connects to related records? Do we know whether it’s appropriate to use? Can we establish provenance, control who – and what – can access it, and trust what comes back?

Those aren’t compliance questions. They’re strategic data questions. Referring the whole subject to legal can close a conversation that should actually be opening.

The decision nobody is making

For decades, historical data was treated primarily as evidence of what had already happened. It was retained because regulations, contracts, policies or operational requirements said it should be. AI changes the economics of that information – because a historical record is no longer valuable only as evidence of the past. Potentially, it can also provide context for future decisions.

That doesn’t mean every historical document should be fed into an AI system. Far from it. It means organisations need to make a deliberate decision about which parts of their historical information estate contain genuinely useful institutional knowledge, whether that information can appropriately be used, and what needs to happen to make it accessible and trustworthy.

That is the historical data strategy I think is missing from many AI programmes.

The CIO I mentioned earlier wasn’t negligent. They were following the structure organisations have used for years: AI treated as a technology investment, historical data treated as a compliance and records-management issue, and nobody connecting the two conversations.

That connection is the plan nobody has made yet.

What making the connection looks like

It doesn’t begin with putting thirty years of documents into an AI platform. It begins with understanding what you actually have. What information exists, where is it, and what condition is it in? Which parts have real value to the AI use cases you’re investing in? What can legitimately be used – and what shouldn’t be? What needs converting, extracting, classifying or enriching? Where has context been lost, what provenance needs to be retained, and what governance has to be in place before an AI system should be allowed anywhere near it?

Only then can you prioritise sensibly. Start with the information most relevant to the business problem the AI is supposed to solve. Make it accessible, structure it appropriately, preserve its context, apply the necessary controls, and build from there.

What that creates is something far more useful than “lots of data.” It creates a governed body of organisational knowledge that AI can retrieve from and use as context – which is what turns the historical archive from something you merely retain into something capable of contributing value.

Where Dajon fits

At Dajon Data Management, this is the natural progression of a journey we’ve been on for nearly thirty years. We started in document storage. Then came digitisation, then data migration and integration. We’re now developing Dajon’s Data Intelligence Solution around the next challenge: how do you turn the historical information an organisation already holds into structured, governed, AI-ready organisational knowledge?

The objective is not simply to digitise more documents. It’s to make the information within them usable – extracted, consistently classified, enriched with meaningful metadata, with the relationships and context that make it valuable preserved – and ultimately to make relevant historical knowledge discoverable by the AI systems organisations are already investing in.

Because for many businesses, the institutional knowledge they want their AI to have isn’t missing. They already own it. They’ve been storing it for decades. The challenge is turning what they have into something their AI can actually use.

The plan nobody has made yet

Organisations are spending extraordinary amounts of money asking what AI can do for them. I think there’s a second question that belongs alongside it: what does our AI need to know about us to do it well?

Then ask where that knowledge currently lives. Some of it will be in modern systems and databases. Some will be inside the heads of experienced employees. A considerable amount may be sitting in the historical information estate the organisation has spent thirty years treating primarily as a cost and a compliance obligation.

That information does not become an AI knowledge base simply because it exists. It has to be selected, made accessible, given context, governed, structured appropriately – and trusted.

Which is why I keep coming back to the same question. Can you really have an AI strategy without a historical data strategy?

If your organisation is investing seriously in AI, what are you doing about the historical data it will need to understand you?