TGDC Contact
SVC-02 / Service TGDC

License operational data that never appeared on the open web

TGDC licenses rights-cleared operational data from real European companies to AI labs: coherent document archives, structured extracts, and held-out evaluation sets. Every corpus arrives with written chain of title and a per-batch anonymisation report.

Formats
PDF · EML · XLSX · DOCX for archives, JSONL · Parquet for extracts
Provenance
Written mandate at the source, documented chain of custody
Per batch
Anonymisation report and checksummed manifest, every delivery

What you can license

Provenance is the product

Every corpus enters through an authorised channel: insolvency estates, which administrators release only under a written mandate, workflow recordings made through consented reenactment, and personal documents contributed by their owners with on-device anonymisation. Nothing is scraped, and no corpus ships without a rights review covering ownership, third-party material and statutory limits.

The result is data with a paper trail: who owned it, who authorised it, what was removed, and what licence terms attach. For labs facing training-data documentation obligations under the EU AI Act, that paper trail is the difference between an asset and a liability.

Anonymisation, verified per batch

Named entities, personal data and commercial identifiers are removed or pseudonymised before delivery. Pseudonymisation is deterministic, so cross-document references stay coherent: the same counterparty maps to the same pseudonym across an entire estate, without exposing who that counterparty was. Each batch ships with a verification report documenting what was detected and replaced.

Why operational data

Web text is exhausted and contaminated: whatever is publicly crawlable is already inside every frontier model. Operational records (how a real company actually corresponded, contracted, invoiced and resolved disputes over years) never appeared online. They carry the longitudinal structure, messiness and domain depth that scraped data cannot provide, which is precisely what document-heavy enterprise AI needs to learn.

How delivery works

Every corpus follows the same four steps: acquired only under written authorisation, cleared through a rights review, anonymised and verified per batch, and delivered as staged corpora with manifests, checksums and licence terms. Licence terms are set per agreement, including exclusive licensing at corpus level, so what you license and on what conditions is documented from the first conversation to the final delivery.

Scope a corpus. Tell us the domain and format you need: hello@thegeneraldata.com. Common questions are answered in the FAQ.