Custom AI training data · India

AI data,
grown, not scraped.

Commission custom training data and evaluations from paid, consented contributors in Indian academic and language communities. Every example comes documented: who made it, under what authorization, for what use.

Early stage · Pre-pilot · Talking to design partners
The garden grows back. Click it.
The problem

The internet's free lunch
is over.

/ 01

Scraped data is a liability

The rights attached to indiscriminately scraped internet data are getting harder to defend. In procurement, in policy, in the press, in court.

/ 02

Bulk data doesn't fix weaknesses

A model failing on Tulu, on Sanskrit, on hard reasoning doesn't need more of the same internet. It needs deliberate, difficult, human examples.

/ 03

Specialist work queues last

Established vendors optimize for scale. A small, unusual job (a low-resource language, a narrow domain, a hard eval) waits behind the volume buyers.

/ 04

The paperwork is becoming the product

Procurement teams, counsel and regulators now ask where training data came from and what may be done with it. A dataset that arrives with its rights, authorizations and lineage answers questions a bare file cannot.

What we make

Commissioned, not collected.

You define the problem: a language, a domain, a failure mode, an eval. We assemble the right cohort, produce the data, and deliver it with its paperwork. Everything below is offered for pilots.

Pilot capability

Custom datasets

Original prompts, responses and domain Q&A written to your spec. Nothing pulled from inventory.

Pilot capability

Evaluation sets

Hard, curated items that probe the exact weakness you're chasing, with accepted answers.

Pilot capability

Red-team scenarios

Adversarial prompts and edge cases designed by humans who understand the failure you fear.

Pilot capability

Preference data

Rankings, critiques and comparisons from screened contributors, adjudicated by reviewers.

Pilot capability

Low-resource languages

Tulu, Sanskrit and code-switched regional text and speech, from communities that actually speak them.

Pilot capability

The paperwork

Dataset cards, consent and compensation records, rights and provenance documentation, QC logs.

Evidence passport — what delivery looks like Specimen · to build
PurposeDefined per pilot: one language, one task, one acceptance test
SourceNewly authored by screened, paid contributors
AuthorizationPer-item consent record; permitted and prohibited uses stated
LicenseBuyer-specific scope: term, territory, exclusivity, in writing
Pay methodAll required time counted; benchmark disclosed
QualityDouble-rated; disputes adjudicated; decisions logged
LineageEvery delivered item mapped to its source records
LimitationsIncluded. What the data is not, stated plainly
The fair chain

Every example has
a paper trail.

01

Contributors opt in, and get paid

Students and language-community members who knowingly participate, compensated for real work. Terms in a language they actually read; a separate yes or no for each use. No disguised labor, no certificate instead of a wage.

02

The task is fixed before work starts

Your problem, acceptance criteria, rights, contributor qualifications and delivery format, agreed up front in writing.

03

Humans produce the data

Contributors create, annotate, evaluate, rank, critique or translate, trained on the task with clear instructions. Unique task variants and process checks mean "human-made" is a control, not a claim.

04

Reviewers keep it honest

Qualified reviewers check the work, settle disagreements, and log every accept and reject. Genuine ambiguity gets reported, not hidden.

05

You receive data with its papers

The dataset, its dataset card, authorization and compensation records, and a provenance trail you can show your lawyers. And your board.

The evidence pack

Documented,
not assumed.

Every pilot delivery ships with six records, so your model team, your counsel and your procurement file all read the same truth. These are the formats the first paid pilot implements.

D-001To build

Dataset card

Purpose, contents, sampling frame, recruiting channels, known limitations. What the data is, and what it isn't.

A-018To build

Authorization record

The contributor's own choices, per use: evaluation, fine-tuning, redistribution. Local-language notice; withdrawal route; voice cloning off by default.

M-001To build

Provenance manifest

Every delivered item mapped to its source record: stable IDs, hashes, transformation history, all versioned.

L-001To build

Rights schedule

Permitted and prohibited uses in writing: term, territory, exclusivity, sublicensing, public release. What your lawyers actually ask for.

Q-001To build

QA report

Calibration results, independent ratings, agreement and adjudication rates, top error categories, method version.

C-001To build

Compensation report

Required time counted, effective hourly rates as distributions, payment timing, deductions, grievances and resolutions.

Illustrative controls, not records of completed client work. Fields, thresholds and formats get agreed with the design partner in the paid pilot. The records exist once the work does.

The network

What we actually have.

No inflated partnerships. Each asset below is stated the way we'd state it across a table, qualifier attached.

Personal connections — not institutional partnerships

Relationships across 10–15 Indian colleges

People we can call. Faculty and departments who know us and can reach motivated undergraduates when a defined pilot lands.

Access — not an employed workforce

Students who want real project experience

Undergraduates seeking compensated, documented work, screened per project and trained per task.

Community access, incl. a traditional gurukul connection

Tulu & Sanskrit speakers and scholars

Rare-language contributors most vendors can't source, for low-resource datasets, evaluation and cultural context.

In the founding circle — structure being formalized

A US-linked company

A related company in the United States gives the venture a commercial bridge to buyers. Its exact role is still being worked out.

Interest expressed — not a customer

An early buyer conversation

One small Silicon Valley AI startup has asked to stay in the loop. A conversation, not a commitment. Anonymous until they say otherwise.

Straight talk

We'd rather tell you
where we are.

This site is a market test and a pilot conversation, not a victory lap. If the proposition lands with you, you're who we built it for.

Founding partner terms
We arePre-pilot, building the operating standard for consent, compensation, QA and provenance.
We areAble to assemble a screened cohort for a defined pilot, through the network above.
We wantOne or two design partners with a narrow, expensive data problem.
Not yetNo datasets delivered, no revenue, no certifications. Not ISO, SOC 2, Fairwork or Fairly Trained. Nothing independently audited.
Not everNo "ethics washing". Clean data doesn't offset anyone's scraped past. It's simply the right way to make new data.
Straight answers

The questions
you'd ask anyway.

Q / 01

You have no customers. Why be first?

Because first is small and specific: a fixed-scope pilot with a calibration gate, milestone payments, written acceptance criteria and a cure period. If we miss, you walk away early, calibration results in hand.

Q / 02

Why not Karya, Appen, Prolific, or an open dataset?

If they meet your task, rights and evidence requirements, use them. Our bet is narrow: founder-led design for hard India language and domain problems, granular contributor authorization, and an evidence pack built for your diligence file. The calibration batch proves whether that's real.

Q / 03

How do you prove a human made the data?

No detector proves that. We combine live onboarding, unique task variants, duplicate and plagiarism checks, spot verification and manual review. Then we report the controls and the residual risk, not a guarantee.

Q / 04

Can a contributor revoke after we train?

The terms distinguish raw data, distributed copies and trained model parameters, and they say so before contribution. We can stop future distribution and pull an item from later dataset versions. Guaranteed "untraining" is not something anyone can honestly promise.

Q / 05

Is your data representative?

Only of its stated sampling frame: recruiting channels, districts, age bands and missing groups, all published with the dataset. A convenience sample doesn't get called representative here.

Q / 06

Does buying this make our model ethical?

No. It gives you one input with evidence behind it, in place of an unverified one. The rest of your training data is still yours to answer for.

Who it's for

Built for the team with
the strange data problem.

AI startups Foundation-model teams Multilingual model developers Evaluation & AI-safety companies Data vendors needing specialist capacity Research laboratories Enterprises with private, domain AI
The pilot

Small, fixed, fast.
Then you decide.

Step 01

Define the problem

One language, domain, or evaluation gap. Narrow on purpose.

Step 02

Fix the scope

Task, acceptance criteria, rights, timeline, and a fixed pilot price.

Step 03

Screen the cohort

Contributors matched to the task; reviewers assigned.

Step 04

Produce & review

Creation with QA, adjudication, and logged decisions.

Step 05

Delivery

Dataset, dataset card, consent and provenance records.

Step 06

Decide

Expand, iterate, or walk away. The data is yours either way.

A well-scoped pilot looks like: one language or domain · ~4 weeks · on the order of 400 authored items / 800 independent judgments · specialist adjudication · versioned dataset + evidence pack · indicative price US$10k–15k, fixed in the SOW before work starts.

Fixed scope & price Calibration gate Milestone payments Written acceptance criteria Cure period Evidence schedule
Your move

Commission the first dataset.

Serious pilot conversations only. We'll bring our questions about your data problem. Bring yours about our process.