AI Data Training
We design and collect datasets and have specialists label them, for fine-tuning, preference training, and evaluation. Our annotators cover finance, health, legal, and code, in more than 30 languages.
We produce training data, run independent model comparisons, set up go-to-market analytics, and run beta tests with real users. Each job ends with a written report you can put in front of a board.
You can hire us for one of these or for all four around a launch. Each one ends with a report or a dataset that belongs to you.
We design and collect datasets and have specialists label them, for fine-tuning, preference training, and evaluation. Our annotators cover finance, health, legal, and code, in more than 30 languages.
We test frontier and open-weight models on your own tasks and report accuracy, latency, cost per conversation, and safety results in one table, with the scripts to rerun it.
We set up event tracking for AI features, from prompt to response to outcome, and build the reports on activation, retention, token cost per user, and pricing.
We recruit testers by country, profile, and device, run the test on TestFlight, Google Play, or the web, and hand back a ranked list of what broke and how to reproduce it.
A study of grading method. We wanted to know how much manual review an LLM judge can take over before its scores stop being reliable. Figures are illustrative.
We took 1,200 real support tickets, removed personal details, and balanced them by intent, language, and length. Two annotators graded every model answer on a five-point scale. A third person settled each disagreement, and that settled grade became the reference.
Once with a single annotator working alone, once with an LLM judge that had the written rubric and two examples, and once with the same LLM judge given no rubric. All three saw the same items in the same order.
Each grader is scored on how often it matched the reference grade. We also report Cohen's kappa, which corrects for the matches you would get by chance on easy items.
We sorted every disagreement by the kind of ticket it came from. That breakdown tells you where a person still has to grade. The overall score does not.
| Grader | Agreement | Kappa | Cost / 1k items | Time / 1k items |
|---|---|---|---|---|
| Two annotators, third settles disagreements | reference | — | $1,900 | 38 h |
| One annotator alone | 91.2% | 0.84 | $780 | 16 h |
| LLM judge with rubric and examples | 88.5% | 0.79 | $14 | 25 min |
| LLM judge with no rubric | 74.0% | 0.58 | $11 | 22 min |