Solutions · LLM training data
Expert data for fine-tuning and aligning LLMs
Prompt–response pairs for supervised fine-tuning, ranked outputs for RLHF, and expert review — including proprietary coding languages and low-resource languages with no public training data.
Overview
Fine-tuning data is expert work, not bulk labeling
Fine-tuning changes how a model behaves, and it learns that behavior from a surprisingly small amount of data. That makes each example count. In supervised fine-tuning (SFT), the model learns from prompts paired with the responses you want. In preference tuning, such as RLHF, it learns from several candidate responses ranked from best to worst. Either way, the data is a set of judgments about what a good answer looks like, and the model copies those judgments — including the inconsistent and wrong ones.
That’s why LLM training data is expert work, not bulk labeling. A ranking is only as good as the person who made it, and fluent model output makes errors easy to miss. The hardest cases are the ones companies care most about: a proprietary coding language that no one outside the company knows, a specialist domain where the right answer takes years of training to recognize, or a language with almost no public text. There’s no crowd to hire for those. The people have to be found, tested, and sometimes taught before they write a single example.
Good fine-tuning data has three things. A written rubric, so every expert applies the same standard. Experts who have proven they can meet it. And review that measures how consistently they did, so you know what you’re training on before you train on it. The output should arrive in your format, with the agreement scores and a record of who wrote and checked each item.
MLtwist builds that data for AI companies. We design the recruiting and skill tests, write and rank the data with sourced experts, run multi-stage review and agreement scoring on every pair, and deliver versioned sets with a Data ID Card. For one AI technology company, that meant training an LLM on a proprietary coding language with annotation timelines more than 40% shorter. If you want your own experts on the hardest cases, they can work alongside ours in the same review process — see how staffing works on the labeling page.
The problem
Why this is hard to do well
No one to hire
A proprietary coding language or a niche domain has no talent pool to recruit from. The people who write and judge the data have to be found, tested, and sometimes taught the language first. A general annotation crowd can't produce data a model should learn from.
Ranking, not labeling
Alignment data isn't one label per item. It's prompts matched to target outputs, and several candidate outputs ranked by correctness and relevance. Every ranking is a judgment call, and the model learns from how consistently those calls are made.
Plausible isn't correct
LLM output reads fluently even when it's wrong, and in code a subtle error can pass a quick look. Reviewers have to know the domain well enough to catch what sounds right but isn't — otherwise the model learns to be confidently wrong.
How it works
How a fine-tuning dataset gets built
Each step is set up so the next one has less to fix: a clear rubric, tested people, and review built into every pair.
- 01
Define the rubric
Agree on the prompt types, the output format, and what makes a response correct, partly correct, or wrong. The rubric is written down before anyone writes data, so every expert applies the same standard and reviewers have something concrete to check against.
- 02
Source and test experts
Recruit coders, linguists, or domain specialists and skill-test them before they touch data. For a proprietary language, the test measures how quickly a candidate can learn it. Only people who pass work on your project, and your own experts can join them.
- 03
Write prompt–response pairs
Experts write prompts matched to the outputs you want the model to produce, in the style, length, and format your model needs. These pairs are the supervised fine-tuning set: examples of the behavior you're teaching, written by people who know what right looks like.
- 04
Rank candidate outputs
For preference data, experts compare several model outputs for the same prompt and rank them by correctness and relevance, from optimal to partially correct to incorrect. The reasoning is recorded, so reviewers can check a ranking instead of guessing why it was made.
- 05
Score agreement
Every pair goes through multi-stage review, and inter-annotator agreement is scored across experts. Disagreements go to a reviewer, not into your set. The scores ship with the delivery, so you can see how consistent the data is before you train on it.
Case study · AI technology company
Training an LLM on a language no one outside the company knows
An AI technology company was building an LLM for a proprietary coding language used in mission-critical applications. The model had to turn natural-language prompts into correct code and rank candidate outputs by quality, and there was no outside talent pool that knew the language.
MLtwist designed a targeted recruitment and testing process to find coders who could learn the language quickly, and ran skill assessments before any annotation began. The coders wrote prompts matched to correct outputs and ranked several code outputs per prompt as optimal, partially correct, or incorrect. Every prompt–code pair went through multi-level review and inter-annotator agreement scoring.
Automation in workflow management and batch QA cut annotation timelines by more than 40%, so the company could iterate on the model faster without raising costs. The model improved at both generating and ranking code in the proprietary language, and the pre-screening meant only top-performing coders contributed to the dataset.
- shorter annotation timelines
- 40%+
What you get
Data your fine-tuning run can use as is
-
SFT pairs
Prompts and target responses in your format, written to an agreed rubric. Ready to load into a supervised fine-tuning run without reformatting.
-
Preference data
Several candidate outputs per prompt, ranked by correctness and relevance, with the reasoning recorded. Ready for RLHF or other preference tuning.
-
Quality numbers
Agreement scores and review results for every pair, not just a final file. You see where experts disagreed and how it was resolved.
-
A Data ID Card
A record of who wrote and reviewed each item, which tools touched it, and under what terms. Useful when someone later asks what the model was trained on.
More work
Related programs
FAQ
Common questions
What's the difference between SFT data and preference data?
Supervised fine-tuning data is prompts paired with the response you want, so the model learns by example. Preference data is several candidate responses to one prompt, ranked from best to worst, so the model learns which of its own outputs to favor. Most alignment programs need both.
Can you work in a language with no public data?
Yes. MLtwist has recruited and tested coders for a proprietary programming language no one outside the company knew, and native speakers for Sindhi, a low-resource language. In both cases, finding and testing the people came first.
Who writes and ranks the data?
Specialists sourced for your project — coders, linguists, or domain experts — who pass a skill test before they start. Your own experts can review samples, take the hardest cases, or work alongside ours. Our team, yours, or both.
How do you measure quality?
Multi-stage review on every pair, plus inter-annotator agreement scoring across experts. Pairs where experts disagree go to a reviewer before they reach your set, and the agreement scores are delivered with the data.
Can you also generate synthetic text?
Yes. Synthetic documents, dialogue, and prompt variants can be generated with LLMs and then reviewed by people. It's useful for widening prompt coverage, but expert-written and expert-ranked data is still what anchors the set.
Do you handle other language data besides LLM training?
Yes. Named entity recognition, relation extraction, coreference, and dependency parsing are all part of past programs — from intelligence analysis on news articles to published research on Sindhi and Global Englishes. See the defense and research pages for those programs.
Talk to us about llm training data
Tell us the data type, volume, and timeline. We'll scope it with our team, yours, or both — and deliver it versioned, in your format.
Also available through Carahsoft and Google Cloud Marketplace.