Rescuing Locked-In HR Records From Legacy Databases

When a FTSE-listed building materials distributor decided to consolidate its employee records onto a single cloud-based HR document platform, the hardest part wasn’t the destination. It was getting the documents out of the systems they were already in.

The client

The client had digitised its personnel files in 2017, prompted by the arrival of GDPR. Records were scanned into an HR document management system, while an older HR application – still in daily use for absence and sickness records – linked out to tens of thousands of files held on a shared network drive.

Both systems had reached the end of their useful life. The main document system couldn’t delete information, not even documents uploaded against the wrong employee, and vendor support had dwindled. As part of a wider HR transformation programme, the client chose to retire both and bring everything together in one platform.

The challenge

On paper, it was a straightforward migration. In practice, the data was effectively locked in.

Every document in the main system was stored inside a 140GB SQL database as a binary large object (BLOB), with no record of file type. Scanned documents weren’t stored as documents at all – each page was a separate object, leaving close to a million individual images, PDFs and Office files to be identified and reassembled. Queries that tried to filter the binary data by file signature ran for extended periods without returning a result.

The network drive held 48,757 files in more than 30 formats, from Word and Excel to bitmaps, Access databases, video and Windows shortcuts. Metadata existed only in filenames, and 4,647 files carried no employee ID at all.

Midway through the project, a further requirement emerged: The older HR application’s absence records also had to move, as they were needed by both HR and Payroll and subject to statutory retention periods. Its absence table held 190,500 rows and 1,440 embedded files, but the application prefixed each file with its own proprietary header, so every extracted document opened as corrupt.

All of this had to be matched to the correct employee, converted to a single standard and delivered ready for upload – without losing anything the client was legally required to retain.

Our approach

Following an on-site discovery workshop with the client’s IT, HR and compliance leads, Dajon took a copy of each database into a dedicated, isolated environment at its London data processing centre, transferred on hardware-encrypted media. Working locally rather than over a remote connection gave the team the processing capacity needed to query and extract the data at speed.

Rebuilding documents from raw binary

Dajon’s developers analysed the database structure and identified the start and end markers written by the scanners used in the original digitisation, along with a hidden index page for every scanned document listing each page in sequence. This allowed nearly a million page images to be reassembled, in the correct order, into complete multi-page documents. Imported files were identified from their file signatures and restored to their original formats.

Stripping hidden headers

For the absence records, the team isolated the proprietary header embedded in each file and removed it to recover the original document. Rather than discarding the header, Dajon retained its upload date and source path as metadata, preserving an audit trail the client would otherwise have lost.

Matching every file to the right employee

Each record was validated against the client’s authoritative employee list, using employee ID as the primary key and accounting for employees whose surnames had changed. Duplicates were flagged by file signature, size and date. Anything that couldn’t be matched, fell outside supported formats, contained macros or was corrupt went into a clearly labelled exception folder, with a full report returned to the client for resolution.

One standard for everything

Every document was converted to a searchable PDF using optical character recognition (OCR), with metadata files mapping each one to an employee and to one of the client’s ten existing HR categories. A Miscellaneous category caught anything that couldn’t be classified, so no record was left behind.

Compliance built in

Alongside the technical work, Dajon shared its HR retention guidance and advised on the compliance implications of the move. This included how right-to-erasure requests interact with statutory retention periods, and the need to safeguard health and safety and payroll-related records that must be kept long after an employee leaves. Dajon also flagged folders earmarked for disposal that contained recent files, recommending a compliance review before anything was discarded.

Throughout, client data was held on hardware-encrypted drives in an environment inaccessible from outside Dajon. Under Dajon’s standard process, all data is erased on sign-off using an HMG IS5 three-pass erasure followed by cryptographic erasure, with a destruction certificate issued to the client.

The results

Nearly 100,000 documents were extracted, converted and prepared for the new platform, with extraction and conversion completed within Dajon’s original estimate of 25–30 working days – despite the complexity uncovered during discovery. Documents were delivered in batches, with Dajon providing batch-level reporting so every upload could be reconciled against the source data. Upload to the new HR platform began in November 2019 and was completed by April 2020, after which the client’s HR team carried out its own quality checks across the migrated records and found no significant issues.

The client now has its employee records in a single, searchable platform, with consistent formats, clear categories and a preserved audit trail – and the legacy systems that had kept its data locked away can finally be retired.