Evidence for AI products

Ground truth for teams building on AI.

We produce training data, run independent model comparisons, set up go-to-market analytics, and run beta tests with real users. Each job ends with a written report you can put in front of a board.

Task accuracy on a support-ticket test set n = 1,200 · illustrative
0 25 50 75 100% Frontier model A Frontier model B Open-weight 70B Current production 84.0 79.0 76.0 71.0
candidate model baseline: current production, 71.0%
Sample chart from a model-selection report. Every model answered the same prompts and was graded by the same people. The scripts that produced the numbers come with the report.
Practices

What we do

You can hire us for one of these or for all four around a launch. Each one ends with a report or a dataset that belongs to you.

01Build

AI Data Training

We design and collect datasets and have specialists label them, for fine-tuning, preference training, and evaluation. Our annotators cover finance, health, legal, and code, in more than 30 languages.

02Choose

LLM Research & Comparisons

We test frontier and open-weight models on your own tasks and report accuracy, latency, cost per conversation, and safety results in one table, with the scripts to rerun it.

03Launch

GTM Data Analytics

We set up event tracking for AI features, from prompt to response to outcome, and build the reports on activation, retention, token cost per user, and pricing.

04Prove

User Beta Testing

We recruit testers by country, profile, and device, run the test on TestFlight, Google Play, or the web, and hand back a ranked list of what broke and how to reproduce it.

Case study

Can an LLM grade model answers as well as a person?

A study of grading method. We wanted to know how much manual review an LLM judge can take over before its scores stop being reliable. Figures are illustrative.

  1. 01

    Build a reference set

    We took 1,200 real support tickets, removed personal details, and balanced them by intent, language, and length. Two annotators graded every model answer on a five-point scale. A third person settled each disagreement, and that settled grade became the reference.

  2. 02

    Grade the same set three more times

    Once with a single annotator working alone, once with an LLM judge that had the written rubric and two examples, and once with the same LLM judge given no rubric. All three saw the same items in the same order.

  3. 03

    Measure agreement with the reference

    Each grader is scored on how often it matched the reference grade. We also report Cohen's kappa, which corrects for the matches you would get by chance on easy items.

  4. 04

    Sort the disagreements by ticket type

    We sorted every disagreement by the kind of ticket it came from. That breakdown tells you where a person still has to grade. The overall score does not.

Agreement with the reference grade n = 1,200 · illustrative
GraderAgreementKappaCost / 1k itemsTime / 1k items
Two annotators, third settles disagreementsreference—$1,90038 h
One annotator alone91.2%0.84$78016 h
LLM judge with rubric and examples88.5%0.79$1425 min
LLM judge with no rubric74.0%0.58$1122 min
Model cost is at list price. Human cost uses an average annotator rate. The full write-up names the judge model, its version, and the date of the run.
Where the LLM judge is reliableChecking format, checking facts against a reference answer, and detecting the language. Agreement was above 95% on these, so they can be automated.
Where it is notTone in complaint tickets, conversations with several turns, and cases where the model refused when it should have answered. Agreement fell to between 71% and 78%, so people still grade these.
What we do nowThe LLM judge grades everything. People grade a 15% sample plus every item the judge flags as uncertain, and a third person settles disagreements. We report the two scores separately.
About

Who we work with

Early-stage startupsTeams that need to pick a model and a data plan before their next funding round.
Product teamsTeams adding AI features to an existing product who need test results before a roadmap decision.
Research groupsLabs that need labelling capacity and outside evaluation without managing vendors themselves.
Regulated companiesTeams in finance, health, and legal, where the rate of unsafe answers decides which model gets used.
Contact

Tell us what you need to decide

We use this only to answer your brief. You will not be added to a mailing list.

Received. If we can help, you will get a written scope.