Insolvency estates as AI training data
When a company is wound down, its complete document estate (years of correspondence, contracts, invoices and ledgers) passes to an insolvency administrator. Under a written mandate from that administrator, the estate can be rights-reviewed, anonymised and licensed as AI training data. This is TGDC's founding sourcing channel.
- Source
- Document estates of wound-down European companies
- Authorisation
- Written administrator mandate before a single file moves
- Typical span
- Decades of correspondence, contracts, ledgers and forms
What a document estate contains
A single wound-down company leaves behind the full paper trail of its operating life: email correspondence with customers and suppliers, signed contracts and their amendments, invoices and dunning letters, ledgers and closing statements, HR files, and the procedural documents of the winding-down itself. Unlike a web scrape, an estate is coherent: the same counterparties, deals and disputes thread through thousands of documents over years, which is exactly the longitudinal structure that document AI needs to learn and that no public dataset has.
The legal path
- Administrator mandate. The insolvency administrator controls the estate's assets, including its data. Licensing proceeds only under a written mandate from the administrator.
- Rights review. Before anything ships, each corpus goes through a review covering ownership, third-party material and statutory limits: what may be licensed at all, and under which conditions.
- Anonymisation. Named entities, personal data and commercial identifiers are removed or deterministically pseudonymised, so cross-document references stay coherent without exposing real parties. Every batch is verified, and a verification report ships with the delivery.
The result is a corpus that is both legal and useful: the administrator's mandate answers the ownership question, the rights review answers the scope question, and the anonymisation report answers the privacy question, in writing, per batch, before anything reaches a buyer.
Why this data matters for AI
Everything publicly crawlable is already inside every frontier model. Estate data never appeared online: it is genuinely unseen, which makes it valuable twice over: as training material that adds signal instead of repetition, and as evaluation and RL-environment substrate where contamination would otherwise invalidate the results. And because it is real operational work rather than curated text, it carries the noise, ambiguity and edge cases that make trained behaviour robust in deployment.
What it looks like as a product
TGDC delivers estates as licensed corpora: archives with documented chain of custody, structured extracts against a customer schema, and held-out evaluation sets. Each delivery includes written chain of title from the administrator mandate onward and the per-batch anonymisation report. Buyers can trace every artefact back to an authorisation, which is what EU AI Act documentation duties increasingly demand.
Administrators and advisors: if you manage estates and want to understand what a data mandate looks like, write to hello@thegeneraldata.com.