We represent the data the frontier AI hasn't trained on.
Security · Precision · Taste
Now licensing
- Reinforcement learning
Cybersecurity RL tasks for LLM training
~3,000 tasksDockerfile / Compose + prompt + grader
- Pre-training
Enterprise private GitHub & CRMs
276+ companiesCode · Comms · Docs · Databases
- Supervised · SFT
Egocentric manipulation demonstrations
10,000 hrs · 20,000 clipsPer-hand labels · 27 industries
How licensing works
Samples first, terms second. Every dataset ships with the rights chain that makes it usable.
- 01
Tell us what you're training
Send a few lines on the model, the use case and rough volume. We reply with representative samples and the full schema.
- 02
Evaluate on your own stack
Run the samples through your pipeline — attach the RL tasks to a training run, score the video labels, profile the code corpus.
- 03
License with a documented rights chain
Commercial terms per buyer. Every dataset ships with its provenance: founder-signed agreements, signed consent, redaction records.
Questions
What buyers ask first
- What does Thea license?
- Proprietary training data that frontier models have not seen: cybersecurity reinforcement-learning environments with automatic graders, egocentric manipulation demonstrations for imitation learning, and full private corpora (code, communications, docs, CRM, databases) from wound-down startups.
- How is the data licensed?
- Under a commercial data license negotiated per buyer. Every dataset has a documented rights chain — signed directly with founders for enterprise corpora, and signed consent with faces and PII redacted for video. We reply to access requests with samples, schema and terms.
- Is any of this scraped from the public internet?
- No. The cybersecurity tasks are handcrafted around novel bugs rather than public 1-days, the video is recorded on real job sites with consent, and the enterprise corpora come from private repositories and systems that were never public.
- How do I evaluate a dataset before licensing?
- Request access and tell us what you are training. We send representative samples and the full schema so you can run your own evaluation before any commitment. Write to hello@theadata.ai.
- Can I contribute a dataset?
- Yes. If you hold proprietary data — private code, CRM, internal communications, documentation or databases — with clean rights, submit it through the site and we will take it from there.