AI data,
grown, not scraped.
Commission custom training data and evaluations from paid, consented contributors in Indian academic and language communities. Every example comes documented: who made it, under what authorization, for what use.
Early stage · Pre-pilot · Talking to design partnersThe internet's free lunch
is over.
Scraped data is a liability
The rights attached to indiscriminately scraped internet data are getting harder to defend. In procurement, in policy, in the press, in court.
Bulk data doesn't fix weaknesses
A model failing on Tulu, on Sanskrit, on hard reasoning doesn't need more of the same internet. It needs deliberate, difficult, human examples.
Specialist work queues last
Established vendors optimize for scale. A small, unusual job (a low-resource language, a narrow domain, a hard eval) waits behind the volume buyers.
The paperwork is becoming the product
Procurement teams, counsel and regulators now ask where training data came from and what may be done with it. A dataset that arrives with its rights, authorizations and lineage answers questions a bare file cannot.
Commissioned, not collected.
You define the problem: a language, a domain, a failure mode, an eval. We assemble the right cohort, produce the data, and deliver it with its paperwork. Everything below is offered for pilots.
Custom datasets
Original prompts, responses and domain Q&A written to your spec. Nothing pulled from inventory.
Evaluation sets
Hard, curated items that probe the exact weakness you're chasing, with accepted answers.
Red-team scenarios
Adversarial prompts and edge cases designed by humans who understand the failure you fear.
Preference data
Rankings, critiques and comparisons from screened contributors, adjudicated by reviewers.
Low-resource languages
Tulu, Sanskrit and code-switched regional text and speech, from communities that actually speak them.
The paperwork
Dataset cards, consent and compensation records, rights and provenance documentation, QC logs.
Every example has
a paper trail.
Contributors opt in, and get paid
Students and language-community members who knowingly participate, compensated for real work. Terms in a language they actually read; a separate yes or no for each use. No disguised labor, no certificate instead of a wage.
The task is fixed before work starts
Your problem, acceptance criteria, rights, contributor qualifications and delivery format, agreed up front in writing.
Humans produce the data
Contributors create, annotate, evaluate, rank, critique or translate, trained on the task with clear instructions. Unique task variants and process checks mean "human-made" is a control, not a claim.
Reviewers keep it honest
Qualified reviewers check the work, settle disagreements, and log every accept and reject. Genuine ambiguity gets reported, not hidden.
You receive data with its papers
The dataset, its dataset card, authorization and compensation records, and a provenance trail you can show your lawyers. And your board.
Documented,
not assumed.
Every pilot delivery ships with six records, so your model team, your counsel and your procurement file all read the same truth. These are the formats the first paid pilot implements.
Dataset card
Purpose, contents, sampling frame, recruiting channels, known limitations. What the data is, and what it isn't.
Authorization record
The contributor's own choices, per use: evaluation, fine-tuning, redistribution. Local-language notice; withdrawal route; voice cloning off by default.
Provenance manifest
Every delivered item mapped to its source record: stable IDs, hashes, transformation history, all versioned.
Rights schedule
Permitted and prohibited uses in writing: term, territory, exclusivity, sublicensing, public release. What your lawyers actually ask for.
QA report
Calibration results, independent ratings, agreement and adjudication rates, top error categories, method version.
Compensation report
Required time counted, effective hourly rates as distributions, payment timing, deductions, grievances and resolutions.
Illustrative controls, not records of completed client work. Fields, thresholds and formats get agreed with the design partner in the paid pilot. The records exist once the work does.
What we actually have.
No inflated partnerships. Each asset below is stated the way we'd state it across a table, qualifier attached.
Relationships across 10–15 Indian colleges
People we can call. Faculty and departments who know us and can reach motivated undergraduates when a defined pilot lands.
Students who want real project experience
Undergraduates seeking compensated, documented work, screened per project and trained per task.
Tulu & Sanskrit speakers and scholars
Rare-language contributors most vendors can't source, for low-resource datasets, evaluation and cultural context.
A US-linked company
A related company in the United States gives the venture a commercial bridge to buyers. Its exact role is still being worked out.
An early buyer conversation
One small Silicon Valley AI startup has asked to stay in the loop. A conversation, not a commitment. Anonymous until they say otherwise.
We'd rather tell you
where we are.
This site is a market test and a pilot conversation, not a victory lap. If the proposition lands with you, you're who we built it for.
Founding partner termsThe questions
you'd ask anyway.
You have no customers. Why be first?
Because first is small and specific: a fixed-scope pilot with a calibration gate, milestone payments, written acceptance criteria and a cure period. If we miss, you walk away early, calibration results in hand.
Why not Karya, Appen, Prolific, or an open dataset?
If they meet your task, rights and evidence requirements, use them. Our bet is narrow: founder-led design for hard India language and domain problems, granular contributor authorization, and an evidence pack built for your diligence file. The calibration batch proves whether that's real.
How do you prove a human made the data?
No detector proves that. We combine live onboarding, unique task variants, duplicate and plagiarism checks, spot verification and manual review. Then we report the controls and the residual risk, not a guarantee.
Can a contributor revoke after we train?
The terms distinguish raw data, distributed copies and trained model parameters, and they say so before contribution. We can stop future distribution and pull an item from later dataset versions. Guaranteed "untraining" is not something anyone can honestly promise.
Is your data representative?
Only of its stated sampling frame: recruiting channels, districts, age bands and missing groups, all published with the dataset. A convenience sample doesn't get called representative here.
Does buying this make our model ethical?
No. It gives you one input with evidence behind it, in place of an unverified one. The rest of your training data is still yours to answer for.
Built for the team with
the strange data problem.
Small, fixed, fast.
Then you decide.
Define the problem
One language, domain, or evaluation gap. Narrow on purpose.
Fix the scope
Task, acceptance criteria, rights, timeline, and a fixed pilot price.
Screen the cohort
Contributors matched to the task; reviewers assigned.
Produce & review
Creation with QA, adjudication, and logged decisions.
Delivery
Dataset, dataset card, consent and provenance records.
Decide
Expand, iterate, or walk away. The data is yours either way.
A well-scoped pilot looks like: one language or domain · ~4 weeks · on the order of 400 authored items / 800 independent judgments · specialist adjudication · versioned dataset + evidence pack · indicative price US$10k–15k, fixed in the SOW before work starts.
Commission the first dataset.
Serious pilot conversations only. We'll bring our questions about your data problem. Bring yours about our process.