License operational data that never appeared on the open web
TGDC licenses rights-cleared operational data from real European companies to AI labs: coherent document archives, structured extracts, and held-out evaluation sets. Every corpus arrives with written chain of title and a per-batch anonymisation report.
- Formats
- PDF · EML · XLSX · DOCX for archives, JSONL · Parquet for extracts
- Provenance
- Written mandate at the source, documented chain of custody
- Per batch
- Anonymisation report and checksummed manifest, every delivery
What you can license
- Archives. Coherent document estates: correspondence, contracts, invoices and ledgers spanning years of real operations, with documented chain of custody. Delivered as PDF, EML, XLSX and DOCX.
- Structured extracts. Entity-resolved, anonymised extractions delivered against your schema, as JSONL or Parquet.
- Evaluation sets. Held-out, never-published document tasks for benchmarking extraction, reasoning and agent workflows on genuinely unseen material.
- RL environments. Task-based training environments with programmatic verifiers, built from the same cleared sources. See RL environments.
- Bespoke corpora. Custom professional datasets assembled on request across healthcare, legal and finance.
Provenance is the product
Every corpus enters through an authorised channel: insolvency estates, which administrators release only under a written mandate, workflow recordings made through consented reenactment, and personal documents contributed by their owners with on-device anonymisation. Nothing is scraped, and no corpus ships without a rights review covering ownership, third-party material and statutory limits.
The result is data with a paper trail: who owned it, who authorised it, what was removed, and what licence terms attach. For labs facing training-data documentation obligations under the EU AI Act, that paper trail is the difference between an asset and a liability.
Anonymisation, verified per batch
Named entities, personal data and commercial identifiers are removed or pseudonymised before delivery. Pseudonymisation is deterministic, so cross-document references stay coherent: the same counterparty maps to the same pseudonym across an entire estate, without exposing who that counterparty was. Each batch ships with a verification report documenting what was detected and replaced.
Why operational data
Web text is exhausted and contaminated: whatever is publicly crawlable is already inside every frontier model. Operational records (how a real company actually corresponded, contracted, invoiced and resolved disputes over years) never appeared online. They carry the longitudinal structure, messiness and domain depth that scraped data cannot provide, which is precisely what document-heavy enterprise AI needs to learn.
How delivery works
Every corpus follows the same four steps: acquired only under written authorisation, cleared through a rights review, anonymised and verified per batch, and delivered as staged corpora with manifests, checksums and licence terms. Licence terms are set per agreement, including exclusive licensing at corpus level, so what you license and on what conditions is documented from the first conversation to the final delivery.
Scope a corpus. Tell us the domain and format you need: hello@thegeneraldata.com. Common questions are answered in the FAQ.