Personal data in licensed training corpora
An archive that a company actually used contains people: names in signature blocks, addresses in invoices, a caseworker in a file note. Whether the corpus a buyer receives still falls under the GDPR is decided before delivery, not after it.
- The rule that decides it
- Recital 26: the GDPR does not apply to anonymous information
- The bar it sets
- Singling out, linkability and inference must all fail
- At The General Data Company
- Gates on every batch, a report with every delivery
Which law reaches a data sale
Four instruments touch the same transaction, and they ask different questions.
- GDPR (Regulation (EU) 2016/679) governs everything that remains personal data. It asks for a legal basis, a purpose, and a record of both.
- AI Act (Regulation (EU) 2024/1689) asks the buyer to document what the model was trained on. Obligations for general-purpose models apply since 2 August 2025, with a published template for the training content summary; the data-governance duties for high-risk systems follow on 2 August 2026.
- Data Act (Regulation (EU) 2023/2854), applicable since 12 September 2025, governs who may use data generated by connected products and services, which decides whether a seller is free to sell at all.
- Data Governance Act (Regulation (EU) 2022/868), applicable since 24 September 2023, sets the frame for intermediaries that broker data between holders and users.
Copyright sits alongside them: Article 4 of the Copyright in the Digital Single Market Directive allows text and data mining unless the rightsholder has reserved the right, and the sui generis database right still protects substantial extractions. That part of the question is covered in provenance under the AI Act.
When a corpus stops being personal data
Recital 26 of the GDPR puts anonymous information outside the regulation entirely. The threshold is not a redaction pass. European guidance asks whether a person can still be singled out in the data, whether records about one person can be linked across the set, and whether an attribute can be inferred about them. A corpus is anonymous when all three fail, taking into account the means reasonably likely to be used by anyone, not only by the buyer.
Pseudonymized data does not meet that bar. Replacing a name with a token leaves the data personal, because the token still points at one person and the pattern around it often identifies them. Treating pseudonymization as anonymization is the most common way a data sale turns into a processing operation nobody documented.
What a delivery carries
- The source authority in writing, from the mandate or consent that allowed the material to leave its origin through to the license the buyer signs.
- Gates before the corpus ships, run on every batch rather than sampled once per supplier, and blocking rather than advisory.
- A report per batch that states what was found and what was removed, in numbers a buyer can check against the delivered files.
The gates are documented by their effect, not by their construction. A buyer receives the counts, the categories and the residue, which is what an auditor asks for. The detection stack behind them is our own work and stays that way, and a description of it would in any case tell a reader nothing about whether a specific delivery is clean.
Where we stand
The General Data Company is GDPR compliant. We process every corpus under European data protection law, from the documented legal basis at the source through to the anonymization that is verified before anything ships, and we hold the records that show it per batch rather than per promise. The same applies to our own operations: the data we hold about the people we work with is handled under the regulation, and requests under it are answered.
One line is worth drawing clearly, because it protects you rather than us. A supplier cannot carry your duties. You train the model, so the obligations that attach to training stay with you, and no license text moves them. Our part is to make them answerable: the legal source is named, the removals are shown, and both arrive in a form that survives an audit.
Four questions before you buy a corpus
- What allowed this material to leave the company that produced it?
- Was the anonymization verified per batch, and can you see the numbers?
- Would the result still be anonymous to someone who holds another dataset about the same people?
- If a supervisory authority writes to you in a year, which document do you send?
If a corpus cannot answer these, the gap does not close by training on it. Ask us the same questions: hello@thegeneraldata.com.