Solutions · Model evaluation & validation
Validation data that holds your model to account
Test sets built to be more accurate than the model they measure, and human review of model output at production scale — scored by class, so you see where a model fails, not just its average.
Overview
A model score is only as good as the test set behind it
Every decision about a model — ship it, retrain it, pick it over another — rests on an evaluation score, and every score rests on a test set. If that set is noisy, too small, or quietly changes between runs, the score is wrong in ways no one can see. Evaluation data is the one place where “good enough” labels aren’t good enough: the labels have to be more accurate than the model they’re judging, or the number reports labeling errors instead of model quality.
A useful evaluation program has four parts. A held-out set that the model never trains on, sampled on purpose to cover the classes, languages, and conditions that matter — especially the rare ones. Labels adjudicated by experts, so disagreements are resolved rather than averaged. Results broken out by class and condition, because a single accuracy number hides exactly the failures you need to find. And versioning, so the next model is scored on the same data and you can tell whether it actually improved.
The second half of evaluation is watching models after they ship. Reviewers look at real predictions and accept, correct, or reject each one, with the reason attached. That turns production output into a measured error record and, if you want, into new training data for the cases the model gets wrong.
MLtwist builds both. Our team, yours, or both label and adjudicate the test set, run review of model output at production volume, and report agreement and miss rates by class. Each set is delivered as a version with a Data ID Card, so you can cite exactly what a model was scored on. For one AI validation platform, that meant reviewing 600,000 audio labels and creating 50,000 more in a month at 98% accuracy — read the Bobidi story. For the labeling and QA process behind it, see data labeling.
The problem
Why this is hard to do well
A test set must beat the model
If the test labels are less accurate than the model being scored, the numbers measure labeling noise, not model quality. A model can look worse than it is, or better. Evaluation data needs a higher bar than training data, and someone has to hold it there.
Errors hide in averages
One accuracy number can look healthy while a model fails on a whole class, a language, or a lighting condition. Those failures usually sit in the rare cases that matter most. You only see them if the test set covers them and the results are broken out.
Scores drift between versions
Each new model version has to be scored on exactly the same data, or you can't tell whether it improved. When test sets are rebuilt, relabeled, or quietly edited between runs, comparisons stop meaning anything, and nobody notices until a regression reaches production.
How it works
How an evaluation set gets built
Evaluation is a small, carefully built dataset plus a repeatable way to score against it.
- 01
Sample for coverage
Pull held-out data that covers the classes, languages, and conditions you care about, including the rare ones that averages hide. The sample is kept apart from anything the model trains on, so a good score can't come from the model having seen the answers.
- 02
Label to a higher bar
Experts label the set and adjudicate every disagreement, because a test set is only as good as its labels. Consensus rules decide how many people see each item. Where your own specialists should make the final call, they review the hard cases.
- 03
Review model output
For models already running, reviewers accept, correct, or reject each prediction and attach the reason. That turns raw model output into a measured error record.
- 04
Measure by class
Agreement and miss rates are reported by class, language, and condition, not only overall. Weak spots show up as numbers you can act on — a class the model confuses, a condition where recall drops — instead of a single score that hides them.
- 05
Version the set
Each test set is saved as a version you can diff and cite. Every model is scored against the same data, so results stay comparable across releases. When the set has to grow, the change is a new version, not an untracked edit.
Case study · AdTech company
Datasets for scoring a brand-safety model
An AdTech company built a product that uses NLP and computer vision to judge whether text, images, video, and audio are safe for ads, following the industry's IAB brand-safety guidelines. It needed structured, high-quality datasets to assess how well its models performed.
MLtwist automated the ingestion, cleaning, and structuring of unstructured content from many sources and produced labeled datasets aligned with the IAB guidelines. The company uses them to measure its models, looking for false positives that block good inventory and false negatives that put ads next to unsafe content.
With reliable data to measure against, the company's models now give more precise contextual brand-safety recommendations, with fewer errors. Less manual annotation cut its data processing costs, and faster model updates let it keep pace with changing content and evolving IAB standards.
Read the case studyWhat you get
A test set you can score every model against
-
A golden set
Held-out data labeled and adjudicated by experts, versioned, and kept apart from training data. Every model you evaluate is measured against the same thing.
-
Review results
Every model prediction marked accepted, corrected, or rejected, with the reason attached. You get an error record, not just a score.
-
Per-class metrics
Agreement and miss rates by class, language, and condition. The weak spots are listed, so you know what to fix before launch.
-
Fast turnaround
Ready-made pipelines and built-in QC do the preprocessing and checks. Evaluation fits into the weeks before launch instead of pushing the date.
More work
Related programs
Bobidi · AI & technology
How Bobidi and MLtwist Delivered Faster AI Validation and Higher Quality at Lower Cost
Retail company · Retail
Retail Company Uses MLtwist for Safe and Accurate Drone Delivery Vision
FAQ
Common questions
How is this different from labeling training data?
The bar is higher and the set is fixed. Test sets are adjudicated by experts, versioned, and kept apart from training data, so every model is scored on the same thing. Training data can tolerate some noise; evaluation data can't, because the noise ends up in the score.
Can you review our model's output at volume?
Yes. For one AI validation platform, MLtwist reviewed 600,000 existing audio labels and created 50,000 new ones in a month, at 98% accuracy against a 95% target.
How big does a test set need to be?
Big enough to cover every class and condition you care about with enough examples to trust the per-class numbers. That usually means sampling rare cases on purpose rather than at random. We scope the size with you based on your classes, languages, and conditions.
Which data types?
Audio, images, video, and text. Past programs include audio label validation, brand-safety classification across text, images, video, and audio, and safety-critical drone vision where recall on people and pets was checked first in QA.
Can our experts make the final call?
Yes. Your specialists can adjudicate the hard cases while our team handles volume — our team, yours, or both. The review process and the record of who decided what stay the same either way.
Can we reuse the set for future models?
That's the point of versioning it. The set stays fixed so new models are compared fairly, and when it needs new classes or conditions, the additions become a new version with its own Data ID Card.
Talk to us about model evaluation & validation
Tell us the data type, volume, and timeline. We'll scope it with our team, yours, or both — and deliver it versioned, in your format.
Also available through Carahsoft and Google Cloud Marketplace.