TGDC Contact
REF-04 / Reference TGDC

Frequently asked questions

The short answers about The General Data Company. Anything missing: hello@thegeneraldata.com.

What TGDC is
An AI training-data company, based in San Francisco, sourcing EU-native
Channels
Insolvency estates, workflow capture, personal documents. Nothing scraped
Contact
hello@thegeneraldata.com

What is The General Data Company?

The General Data Company (TGDC) is an AI training-data company. It sources real operational data from European companies, anonymises it, and licenses it to AI labs for training, evaluation and product development. TGDC is based in San Francisco, USA; data sourcing is EU-native and GDPR-cleared.

What does TGDC sell?

Four things: document archives with documented chain of custody, structured extracts delivered against a customer schema, held-out evaluation sets built from never-published material, and RL environments (task-based training environments with programmatic verifiers built from real workflows). Bespoke corpora are assembled on request.

What are RL environments?

An RL environment packages a task specification, realistic working state, and a programmatic verifier that scores each attempt, so a language model can be trained with reinforcement learning on real work. TGDC builds them from consented, rights-cleared enterprise workflows that never appeared on the open web.

Where does the data come from?

Three authorised channels: insolvency estates, which administrators release only under a written mandate, workflow capture recorded through consented reenactment, and personal documents contributed by individuals through a contributor app, not yet released, with on-device anonymisation. Nothing is scraped.

Is the data GDPR-compliant?

Sourcing is EU-native and GDPR-cleared. Every corpus passes a rights review covering ownership, third-party material and statutory limits, and personal data is removed or pseudonymised before delivery, verified per batch.

How is the data anonymised?

Named entities, personal data and commercial identifiers are removed or deterministically pseudonymised: the same real-world party maps to the same pseudonym across a corpus, so documents stay coherent without exposing anyone. Each batch ships with a verification report.

Does TGDC scrape the web?

No. Every corpus is authorised at the source and carries a written chain of title. That is the point: the data TGDC licenses never appeared on the open web, so it is genuinely unseen by existing models.

Can we license data or environments exclusively?

Yes. Exclusive licensing is available at the task-collection and corpus level, delivered to one customer and withheld from every other. Terms are set per agreement; write to hello@thegeneraldata.com.

Is The General Data Company an IT services or label-printing company?

No. The General Data Company at thegeneraldata.com is an AI training-data company and is not affiliated with any similarly named IT services, computer networking or label-manufacturing business.

What formats do deliveries use?

Archives ship as PDF, EML, XLSX and DOCX with manifests and checksums; structured extracts as JSONL or Parquet against your schema. Every delivery includes provenance documentation, the per-batch anonymisation report and licence terms.