The General Data CompanyContact
REF-05 / PrimerThe General Data Company

Personal data in licensed training corpora

An archive that a company actually used contains people: names in signature blocks, addresses in invoices, a caseworker in a file note. Whether the corpus a buyer receives still falls under the GDPR is decided before delivery, not after it.

The rule that decides it
Recital 26: the GDPR does not apply to anonymous information
The bar it sets
Singling out, linkability and inference must all fail
At The General Data Company
Gates on every batch, a report with every delivery

Which law reaches a data sale

Four instruments touch the same transaction, and they ask different questions.

Copyright sits alongside them: Article 4 of the Copyright in the Digital Single Market Directive allows text and data mining unless the rightsholder has reserved the right, and the sui generis database right still protects substantial extractions. That part of the question is covered in provenance under the AI Act.

When a corpus stops being personal data

Recital 26 of the GDPR puts anonymous information outside the regulation entirely. The threshold is not a redaction pass. European guidance asks whether a person can still be singled out in the data, whether records about one person can be linked across the set, and whether an attribute can be inferred about them. A corpus is anonymous when all three fail, taking into account the means reasonably likely to be used by anyone, not only by the buyer.

Pseudonymized data does not meet that bar. Replacing a name with a token leaves the data personal, because the token still points at one person and the pattern around it often identifies them. Treating pseudonymization as anonymization is the most common way a data sale turns into a processing operation nobody documented.

What a delivery carries

The gates are documented by their effect, not by their construction. A buyer receives the counts, the categories and the residue, which is what an auditor asks for. The detection stack behind them is our own work and stays that way, and a description of it would in any case tell a reader nothing about whether a specific delivery is clean.

Where we stand

The General Data Company is GDPR compliant. We process every corpus under European data protection law, from the documented legal basis at the source through to the anonymization that is verified before anything ships, and we hold the records that show it per batch rather than per promise. The same applies to our own operations: the data we hold about the people we work with is handled under the regulation, and requests under it are answered.

One line is worth drawing clearly, because it protects you rather than us. A supplier cannot carry your duties. You train the model, so the obligations that attach to training stay with you, and no license text moves them. Our part is to make them answerable: the legal source is named, the removals are shown, and both arrive in a form that survives an audit.

Four questions before you buy a corpus

  1. What allowed this material to leave the company that produced it?
  2. Was the anonymization verified per batch, and can you see the numbers?
  3. Would the result still be anonymous to someone who holds another dataset about the same people?
  4. If a supervisory authority writes to you in a year, which document do you send?

If a corpus cannot answer these, the gap does not close by training on it. Ask us the same questions: hello@thegeneraldata.com.