Data Engineering

From Raw Dumps to
Mission-Ready Records

Getting hold of hacked, breached and leaked data is only half the problem. Shadow Nexus repairs, normalizes, translates, validates and entity-resolves breach and leak data so analysts query records instead of wrestling with files — against our holdings, or against yours.


The problem with raw breach data

A leaked database does not arrive as a tidy table. It arrives as a multi-gigabyte archive with mixed character encodings, inconsistent delimiters, embedded line breaks inside quoted fields, truncated SQL that no client will import, column headers in Mandarin or Farsi or Spanish, duplicate records recirculated from an older dump under a new name, and sometimes no real data at all — just a convincing fabrication built to be sold.

Every one of those problems has to be solved before a single analytic question can be asked. Most organizations discover this after the data is already in-house, and the dataset quietly becomes shelfware.

The pipeline

Acquire

Datasets come from existing Shadow Nexus holdings, from tasked collection against a stated requirement, or from material the customer already owns and cannot make usable.

Repair

Character encoding detection and correction, delimiter and quoting repair, recovery of rows that standard parsers reject, and reconstruction of truncated or malformed SQL. Failed rows are worked through iteratively rather than discarded, because the broken rows are often the interesting ones.

Normalize and translate

Foreign-language and non-obvious column headers are mapped to a single common schema, so a national identifier field means the same thing whether it arrived labelled in Chinese, Arabic, Russian or Spanish. Dates, phone formats, name order and address structures are standardized per locale rather than force-fitted to a US shape.

Validate authenticity

Each dataset is assessed on intrinsic structural evidence before it is trusted: national identifier checksum decoding, value distribution analysis, sequential identifier detection, temporal flatness testing and cleartext secret inspection. Fabricated and synthetic datasets are identified and rejected.

De-duplicate against holdings

Incoming material is matched against existing holdings on whole-record field composition, not just identifier overlap, so a repackaged dump is recognized as something you already have rather than bought and ingested a second time.

Resolve entities

Records are linked across datasets, languages and document types into consolidated person and organization views, with the source of every attribute preserved so an analyst can defend a finding back to the record that produced it.

Deliver

Output lands where the mission runs: your analytics platform, a search index, a columnar store, or the OcientAIQ Unified Data Platform, in cloud, on-premises, hybrid, tactical or air-gapped environments.

Authenticity validation

Fabricated datasets are a real and growing part of the market. They are cheap to generate, superficially convincing, and extremely expensive to discover after an agency has built a finding on top of one.

Shadow Nexus treats authenticity as a gate, not a disclaimer. Validation runs on the structure of the data itself rather than on the seller’s claims, which means it works even when a dataset has no public counterpart to compare against — the normal case for genuinely closed-source material.

  • Checksum decoding — national identifiers that carry internal check digits, embedded birth dates or region codes are decoded and tested for internal consistency.
  • Distribution analysis — real populations produce characteristic distributions of names, ages, regions and issue dates; generated data rarely does.
  • Sequential and temporal testing — artificial identifier runs and unnaturally flat timestamp patterns are strong fabrication signals.
  • Cross-reference against known-real holdings — where overlap exists, records are corroborated or contradicted individually rather than judged in aggregate.

Entity resolution

A person does not appear the same way twice across unrelated breaches. Names transliterate differently, phone numbers carry different prefixes, national identifiers are formatted inconsistently and addresses are recorded in whatever structure the originating system happened to use.

Shadow Nexus resolves those fragments into consolidated views — person, organization, and the links between them — so an analyst asks about a target rather than about a table. The result is attribution-grade: every attribute on a resolved entity points back to the dataset and record it came from.

Where the output lands

Your platform

Delivered in the schema and format your existing stack expects, ready to ingest without further engineering.

Search index

Indexed for selector-driven investigation, with parent and child record relationships preserved.

Columnar store

Loaded for analytical workloads over large record volumes where scan speed matters more than lookup.

OcientAIQ

Ingested into the OcientAIQ Unified Data Platform for full-fidelity analysis alongside your own mission holdings.

Two ways to engage

Our data

Processed holdings

Take Shadow Nexus HBL datasets already repaired, validated and entity-resolved. Nothing to build, nothing to clean.

Your data

Engineering as a service

Hand over material you already hold and cannot use — acquired, inherited or seized — and get it back repaired, normalized, validated, de-duplicated and entity-resolved, in your environment under your control.

Frequently asked questions

Can Shadow Nexus process data we already hold?

Yes. The same pipeline runs against customer-supplied material, including datasets that have sat unusable for years. Where the work has to stay inside your environment, it can be run there rather than shipped out.

What formats can you work with?

Delimited text in any encoding, SQL dumps including truncated and malformed ones, JSON and newline-delimited JSON, spreadsheets, fixed-width exports, and the mixed archives that leaked data actually arrives in rather than the tidy formats vendors assume.

Do you handle non-English data?

Yes, and it is the normal case. Column headers and record values are mapped from the source language into a common schema, with locale-correct handling of name order, dates, addresses and national identifier formats rather than a blind English transliteration.

How do you prove a dataset is genuine?

Through intrinsic structural evidence — checksum decoding, distribution analysis, sequential identifier detection and temporal testing — plus per-record cross-reference against known-real holdings where overlap exists. The assessment is reported per dataset, not asserted.

Will you tell us if we already have the data?

Yes. De-duplication runs on whole-record field composition against existing holdings, so a repackaged dataset is flagged as a duplicate before anyone pays for it twice.

Can this run in a classified or disconnected environment?

Yes. The pipeline can be delivered for execution inside an air-gapped or on-premises environment, with no requirement for the data to leave customer control.

Send us the file nobody could open

Describe the material and the environment it has to run in, and Shadow Nexus will tell you what is recoverable, whether it is genuine, and what it looks like once it is resolved.

Contact Shadow Nexus