The Data Fix Your AI Needs Is Smaller Than You Think

Most technology leaders have already had the realisation. The AI is deployed, the outputs are underwhelming, and after a couple of honest diagnostics it becomes clear the constraint isn’t the model – it’s what the model has to work with. That part of the argument is largely settled now.

The harder moment comes next. Someone at the board table says “so let’s fix the data,” and the honest answer is a programme that runs for eighteen months, costs more than the AI investment it’s meant to rescue, and produces nothing visible until year two. At that point the conversation usually stalls, and the organisation quietly returns to tuning the model – not because anyone thinks that will work, but because it’s the only option that fits in a quarter.

That stall is the real problem, and it rests on a premise worth challenging: that making data AI-ready means bringing the whole estate to a uniform standard first.

The enterprise-wide clean-up is the route that fails

Gartner expects 80% of data and analytics governance initiatives to fail by 2027, and its analysts are blunt about the reason – a programme that doesn’t enable prioritised business outcomes fails, and the recommendation is to rescope governance around tangible outcomes rather than broad coverage[1].

Anyone who has watched one of these programmes from the inside recognises the pattern. Scope expands because every domain has a legitimate claim on it. Momentum depends on executive attention that gets redirected after the first reorganisation. The deliverable is a framework rather than a working capability, and by the time it lands the use case that justified it has moved on.

“AI-ready” isn’t a state your data can be in

The deeper issue is that the premise is wrong on its own terms. Gartner’s position is that there is no way to make data AI-ready in general or in advance, because readiness depends entirely on how the data will be used – and that high-quality data, judged by traditional data-quality standards, is not the same thing as AI-ready data[2].

That last point catches people out. Conventional data quality work strips outliers and smooths inconsistencies so the numbers behave for human readers. An AI use case often needs precisely the opposite – the edge cases, the odd claims, the exceptions, because those are the situations it will actually encounter. Data can be immaculate by the standards of a governance scorecard and still be the wrong data for the decision you’re trying to support.

Which means the question a CIO should be answering isn’t “is our data ready.” It’s narrower and considerably more answerable: which decisions are we asking AI to support, and what would each of those decisions need to be able to read?

Why the distinction between training and retrieval matters here

There’s a conflation running through most discussion of this topic that quietly makes the problem look bigger than it is. People talk about AI “learning from” their organisation’s documents, which implies the corpus has to be assembled and cleaned wholesale before the model is trained on it. For the great majority of enterprise deployments, that isn’t what’s happening.

The model arrives already trained. What it needs from you is the ability to find and read the right information at the point a question is asked – a retrieval problem rather than a training one. The practical consequence is significant: retrieval can be built one decision at a time. Making the claims files for a particular product line readable doesn’t require the policy archive from a different division to be in the same condition. The work partitions in a way an enterprise data-quality programme does not.

This is also why performance can improve without touching the model, the platform or the architecture. You’re not rebuilding the system. You’re widening what it can see.

What scoped work actually looks like

Start from a decision that matters and has a measurable outcome – a fraud referral, a renewal recommendation, a first-line compliance check. Establish what a competent human would need to read to make that call well, and how much of it currently sits in documents a system cannot meaningfully parse. That gap, for that decision, is the piece of work. It is usually a fraction of the estate, and it delivers something demonstrable inside a quarter rather than a framework inside two years.

Dajon’s Data Intelligence Solution is built for that scoped approach: reading the documents relevant to a given decision, extracting and standardising the entities inside them, and connecting each to the policies, claims and cases it belongs to. It runs alongside existing platforms rather than replacing them, so the structured output feeds the systems already in place.

That last point tends to matter more to technology leaders than anything else. A CTO at an insurance firm who saw the approach recently made it the thing he came back to – that it could sit as a layer over an existing PAS or CAS rather than arriving as another implementation to sequence and integrate. His interest wasn’t in the capability in the abstract. It was that it could be added to what he already ran without a migration.

The question worth taking into the next planning cycle

The instinct when data is identified as the constraint is to size the whole problem. That instinct is what produces the eighteen-month estimate, the stalled conversation and, on Gartner’s numbers, the four-in-five chance of failure.

The more productive question is which single decision, made better, would be worth the most this year – and what that one decision needs to be able to read. Answer that, deliver it, and you have both a working capability and the evidence to fund the next one. The data problem is real. It does not have to be approached all at once.

Dajon’s Data Intelligence Solution extends the knowledge base of existing AI systems — transforming unstructured organisational data into structured, AI-ready assets that improve performance without replacing the technology investment already made. Get in touch to understand where your current data environment might be limiting what your AI can deliver.


References

  1. Gartner Predicts 80% of D&A Governance Initiatives Will Fail by 2027 Gartner[]
  2. AI-Ready Data Essentials to Capture AI Value Gartner[]