Training-data provenance under the EU AI Act
Provenance is the documented answer to two questions: where did this training data come from, and what right do you have to use it? Under the EU AI Act, whose obligations have been phasing in since 2025, AI providers must be able to answer both, in writing.
- The obligation
- A sufficiently detailed summary of the content used for training
- Scraped data
- No chain of title, so no summary anyone can stand behind
- At TGDC
- Rights review per corpus, anonymisation verified per batch
What the Act expects
Providers of general-purpose AI models must maintain technical documentation about their training data and publish a sufficiently detailed summary of the content used for training, alongside a policy for respecting EU copyright law, including reservations of rights expressed by rightsholders. Providers of high-risk AI systems face data-governance duties over the datasets they train on. The direction across all of it is the same: "we crawled it" stops being an answer. Regulators, customers and courts increasingly expect a documented chain from each dataset back to an authorisation.
Scraped versus licensed
- Scraped data has no chain of title. Its copyright status is contested, opt-outs are hard to honour retroactively, and its contents are, by construction, already inside every competing model.
- Licensed data arrives with a written answer to the provenance question: who owned it, who authorised its use, what was paid, and what terms attach. Documentation that regulation asks for is simply part of the delivery.
What rights-cleared means at TGDC
Every corpus The General Data Company delivers carries its provenance with it:
- Written chain of title from the source authorisation (an insolvency administrator's mandate, a consented workflow reenactment, or an individual's in-app consent) through to the licence the buyer signs.
- Rights review per corpus covering ownership, third-party material and statutory limits before anything ships.
- Per-batch anonymisation reports documenting that named entities, personal data and commercial identifiers were removed or pseudonymised, and verified.
A lab that trains on such a corpus can put a documented sourcing story into its technical documentation instead of a description of a crawl. When an auditor or a customer later asks where a dataset came from, the answer is a written chain of title rather than a best guess. See data licensing for the product detail, or insolvency estates for how the founding channel works.
A short provenance checklist for AI teams
- Can you name the legal source of each training dataset: a licence, a mandate, a consent?
- Would the sourcing story survive being written into your model's technical documentation?
- Do your datasets ship with anonymisation or PII-handling evidence, per batch rather than per vendor promise?
- If a rightsholder or regulator asks tomorrow, how long would the answer take?
None of these questions get easier after training has finished. If any answer is uncomfortable, that is the gap rights-cleared licensing closes: hello@thegeneraldata.com.