

The Data Goldmine
On the brink of AGI, a strange new industry is printing money. Twelve-person teams are closing $30M quarters selling training data to frontier labs, and the product depreciates every time a model gets smarter.
We're spending an evening pulling that world apart:
Pipelines: how frontier training data actually gets manufactured: expert task generation, RL environments with rubrics and verifiers, QA through consensus grading and camera-verified provenance, dedup and decontamination before anything ships to a lab
What labs buy right now: post-training expert judgment (SFT and preference data from doctors, physicists, lawyers), simulated enterprise environments for agent RL, physical capture from sensor rigs, and why plain human labeling is already dead
Economics: purchase orders instead of ARR, tasks discarded once models pass them 70% of the time, 4-5x exclusivity multiples, gross versus net once the experts get paid, and who's left standing when the gold learns to mine itself
Talks, a panel, then food and arguing.
Hosted at AGI House.