Our mission

Language models have run out of internet. The next dataset is the physical world.

Language models learned from text anyone could scrape. Robots need something nobody has at scale yet: structured, verified footage of skilled people doing real tasks with their hands. That data doesn't exist in a warehouse waiting to be labeled. It has to be recorded, by the people who actually know how to do the work.

Sourceify exists to make that recording worth doing. Enterprises post exactly what they need, and experts (mechanics, technicians, electricians, lab staff) record themselves doing it and get paid per approved submission. No hiring pipeline, no data-labeling farm, no staged studio shoot.

Why this, why now

Every dataset era needs a new source.

Already tapped

Text

Language models learned from web-scale text that already existed. That well is largely dry.

The opportunity

Physical tasks

Robotics needs footage of skilled manual work that mostly doesn't exist anywhere, recorded on request.

The hard question

Yes, this is training data for automation. We think that's exactly why it has to be built this way.

Physical automation is coming whether or not Sourceify exists. That training data gets collected somewhere: scraped from public footage with no consent asked, contracted from a studio that pays a flat day rate and keeps every right to what it films, or bought from a labeling farm several steps removed from anyone who's actually done the work.

We don't think that path makes the underlying shift any smaller. But we do think there's a real difference between it and one where the person whose skill is being captured chooses to record it, is told exactly what it's for, and gets paid every single time their specific submission is used. Not a one-time buyout, not a work-for-hire clause buried in a contract.

That doesn't resolve the tension. It means the people whose knowledge is the actual raw material get some of the value created from it, and a say in whether it happens at all, instead of none either way.

Priced per task

Every task sets its own price up front — paid directly to the expert who recorded it, no middleman.

Opt-in, every time

See the task, know what it's for, then record — or walk away.

Also worth asking directly

Isn't this just another AI company harvesting people's data?

Short answer: no. The longer answer is in how the money actually moves. We don't collect anything about you in the background. No location, no browsing history, no profile that builds up over time. The only thing we ever get is footage of one task, and you chose to record it, after seeing what it was for and what it paid.

That footage doesn't sit around unedited, either. Faces and anything readable in frame, a screen, a piece of mail, an ID card, get blurred automatically before the recording even saves to your phone. Nobody reviews it after the fact and decides what to blur. It happens live, every time, during the recording itself.

We do take a cut, disclosed to the enterprises who pay it, not hidden from anyone. But it's a fee on work you got paid to produce, not a resale of who you are. You're paid per submission, every time one gets approved, not once upfront for open-ended rights to your face or your work afterward.

That's the actual mechanism, not a slogan on a page.

In practice, not just in principle

Here's what that actually looks like, today.

01

Faces are blurred before the video ever leaves your phone

Face and readable-text detection runs live, frame by frame, during recording itself. It is not a server-side cleanup step after upload.

02

You see the task and the consent terms before you record

Every task shows its instructions and what the footage is for up front. Nothing records until you've agreed.

03

Payment is per submission, not a flat buyout

Every individual approved recording pays out on its own. There's no single lump sum that covers unlimited future use of your likeness or work.

04

Raw enterprise access has a hard expiration

An enterprise's direct access to a campaign's raw data is time-boxed. 30 days after they first pull it, it's gone, not held indefinitely.

What actually happens

From a phone on someone's head to a line in a training set.

01

Recorded

An expert does the task exactly as they normally would, phone in hand or head-mounted. No script, no retakes for the camera.

02

Validated

Automated bounding-box detection and an AI reviewer confirm the right task is happening and hands stay in frame before anything counts as approved.

03

Structured

Video, motion telemetry, and metadata are packaged together, not a raw file dump waiting to be labeled after the fact.

04

Trained on

Enterprises pull approved packages straight into their training pipeline, one task campaign at a time.

What we believe

01

Pay for the work, not the résumé

The person who's replaced a hundred brake pads is the best source of data on replacing brake pads. No application, no interview, no portfolio required.

02

Data that reflects reality

Real hands, real tools, real mistakes and corrections, not a staged studio take. That texture is what a model actually needs, and a script can't fake it.

03

Consent and control stay with the expert

You see the task before you record it, you're paid per submission you choose to make, and your personal information never ships with the data.

04

The value flows to the person who created it

Every approved submission pays the expert who recorded it directly. No studio, no staffing agency, no scraped-dataset broker sitting in between.

05

Open to anyone who can do the work

No degree, no portfolio, no prior "data" experience. If you can do the task in front of a camera, that's the whole bar.

The physical world is the next frontier for robotics.

Help train it, or come collect the data you need.