Destruction Is an AI Readiness Activity

For most of my career, deletion has been the decision nobody wanted to make. Keeping information felt safe. Destroying it felt like a risk: the one file you shredded would turn out to be the one a regulator, a litigant or a client asked for. So organisations kept everything, and the archive grew year after year.

AI changes that calculation. When an AI system can read your information estate, what you keep is no longer just a storage cost or a compliance question. It becomes the material your AI learns from and the source it quotes when it answers your staff and customers. So deciding what to destroy is now part of getting ready for AI.

Keeping everything felt like the safe option

The scale of over-retention has been visible for years. Veritas’s Global Databerg Report, a survey of 2,550 senior IT decision-makers across 22 countries, found that 52% of the information organisations store is “dark” data whose value is unknown, while a further 33% is redundant, obsolete or trivial (ROT)[1]. IT leaders considered just 15% of stored data to be business-critical[1].

That research is now a decade old. Nothing I’ve seen since suggests the underlying habit has changed. What has changed is that we’re now asking AI to make sense of it all.

AI has no instinct for what’s out of date

Experienced people learn to work around a cluttered archive. A claims handler knows the 2011 procedures manual was replaced twice. A paralegal knows which draft of a contract was signed. A relationship manager knows which of four client summaries is the current one. None of this is written down. It’s professional judgement built up over years.

AI doesn’t have that judgement unless the information estate provides it. Many retrieval systems find content because it is relevant to the question, not because it is current, authoritative or still correct. A superseded policy that closely matches the wording of a question can be found just as easily as the policy that replaced it. The AI then presents it with the same confidence.

This isn’t a theoretical concern raised by records managers. Microsoft warned that oversharing in Copilot deployments can lead to outdated or irrelevant responses from AI, undermining its utility[2]. Its own deployment guidance goes further and advises organisations to apply retention and deletion policies to remove inactive or obsolete files to improve the quality of Copilot responses[3].

It’s worth pausing on that. One of the largest AI vendors in the world is telling its customers that deleting things will make its AI work better.

Over-retention was already a compliance problem

None of this should surprise anyone in information governance. In the UK, the storage limitation principle has always pointed the same way. The ICO’s guidance is clear that you must not keep personal data for longer than you need it, and you should periodically review the data you hold, and erase or anonymise it when you no longer need it[4].

What AI adds is exposure. Personal data kept past its retention period used to sit unseen in a file share or a box in a warehouse. The risk was real but mostly theoretical, because nobody went looking. An AI assistant indexing that file share doesn’t need anyone to go looking. It may surface the information on its own, in answer to a question from someone who never knew it existed.

Over-retention used to be a hidden liability. With AI connected to the estate, it can turn up in front of your staff.

Defensible, not indiscriminate

To be clear, I’m not arguing for a clear-out. Organisations hold a great deal of historical information that is genuinely valuable, and much of my argument elsewhere is about unlocking it. Indiscriminate destruction is as much a governance failure as indiscriminate retention.

The point is that every item in the estate needs a decision, and there are really only three:

  • Keep and make usable. This is information with lasting value: precedent, decisions, contracts, case history. It should be classified, given context and made accessible to AI with the right permissions.
  • Keep but ring-fence. This is information you must retain for legal or regulatory reasons but which has little value as AI context. It should be preserved and protected, and kept out of what the AI can see.
  • Destroy. This is information past its retention period that no legal hold or business need applies to. It should be disposed of securely, with an audit trail that shows it was done properly.

Most organisations have only ever properly run the second category. AI readiness means doing all three deliberately.

Don’t pay to digitise what you should destroy

There is a practical consequence that many AI programmes miss. When organisations decide to bring historical paper records into their AI strategy, the natural first step is digitisation. If that happens before a retention review, the organisation pays to convert its ROT into searchable, machine-readable ROT. The clutter that used to be safely out of reach in a box becomes clutter the AI can read.

Sorting before scanning is not a new idea in records management. AI simply makes skipping that step much more costly.

Where this belongs in the AI programme

Gartner predicts that through 2026, organisations will abandon 60% of AI projects unsupported by AI-ready data[5]. Most discussion of AI-ready data is about adding things: metadata, structure, integration, quality controls. Far less of it is about taking things away.

If I were building an AI readiness plan today, I would want to see five things:

  • Disposition decisions treated as part of AI readiness, not left as a separate records-management task.
  • Retention schedules applied before content is indexed or connected to AI tools, not afterwards.
  • Superseded versions of policies, procedures and templates identified and either removed or clearly marked.
  • Retention reviews carried out before any bulk digitisation of paper archives.
  • Disposal treated as an ongoing discipline, because today’s documents are tomorrow’s historical data.

How Dajon approaches it

At Dajon, we’ve spent nearly thirty years on every stage of the information lifecycle: storing physical records, digitising them, managing retention and securely destroying what’s no longer needed. That combination means we see where things go wrong and in what order. We help organisations decide what is worth making AI-ready, what should be protected and kept out of reach, and what should be destroyed securely and defensibly before it causes a problem.

If you’re preparing your historical information for AI, the first question may not be “what can AI do with this?” but “should we still have this at all?”

Can you really have an AI strategy without a historical data strategy? And can you have a historical data strategy that never decides what to let go?


References

  1. Veritas Global Databerg Report Finds 85% of Stored Data Is Either Dark, or Redundant, Obsolete, or Trivial Veritas[][]
  2. From Oversharing to Optimization: Deploying Microsoft 365 Copilot with Confidence Microsoft Community Hub[]
  3. Configure a Secure and Governed Foundation for Microsoft Copilot Microsoft Learn[]
  4. Principle (e): Storage Limitation ICO[]
  5. Lack of AI-Ready Data Puts AI Projects at Risk Gartner[]