Solutions · Data provenance & compliance
Know where every training file came from
Every dataset ships with a Data ID Card: where each file came from, what changed it, who labeled it, and under what terms. Recorded as the work happens, not reconstructed for an audit.
Overview
A training set is only as defensible as its paper trail
Provenance is the history of a piece of data: where it came from, under what terms, what changed it, and who handled it along the way. For a training set, that history used to be an afterthought. Now it’s a question that buyers, auditors, procurement reviewers, and research peer reviewers ask directly — what was this model trained on, and did you have the right to use it? A folder of contracts and a vendor’s word don’t answer it.
The trouble is that training data passes through many hands. A file might be captured by one contractor, cleaned by a script, pre-labeled by a model, annotated in a third-party tool, and reviewed by a different team. Each hand-off is a place where the record breaks. Good provenance closes those gaps: the history is attached to each file, it’s written as the work happens rather than reconstructed later, and it follows the data through every version you deliver.
MLtwist records it in a Data ID Card. When a file enters the platform, its origin, owner, and usage terms are captured. Each Twist that transforms or pre-labels it adds an entry, including the model and prompt behind any draft label. Labelers, reviewers, vendors, and tools are recorded along with how each met your requirements. Every delivery is a version you can diff, roll back, and cite. Stanford researchers used the Data ID Card to audit their data’s path for two published studies, and regulated programs rely on the same record for chain of custody.
Use it when you need to show where your training data came from — for a procurement review, a customer’s security questionnaire, a published paper, or your own confidence before a model ships. It comes with every MLtwist delivery, whether the data was collected, generated, or brought by you, and it’s central to our public sector work.
The problem
Why this is hard to do well
Regulators are asking
AI rules and procurement reviews increasingly ask what a model was trained on and whether you had the right to use it. Answering with a spreadsheet assembled after the fact doesn't hold up. The record has to exist before the question does.
Long vendor chains
Data passes through collectors, tools, and labeling vendors, and each hand-off loses the paper trail. By the time a file reaches training, nobody can say who touched it or what changed. One weak link undoes the record for the whole dataset.
Rights and consent
Licenses, consent, and usage limits have to follow each file, not sit in a contract folder. A clip recorded under one agreement can end up in a dataset used for something else. Without per-file terms, you can't tell which data you're allowed to use for what.
How it works
How a Data ID Card gets filled in
The platform writes the record as the work happens. No one pieces it together later.
- 01
Record the origin
When a file enters MLtwist, its source, owner, capture details, and usage terms are recorded, including consent for collected data. For data MLtwist collects, that means capture location, ownership, associated contracts, and usage rights are on file from the moment the recording is uploaded.
- 02
Log every change
Each Twist that cleans, converts, or pre-labels the file is logged with its version and settings, including the model and prompt behind any pre-label. The log comes from the pipeline itself, so it can't drift from what actually happened to the data.
- 03
Record the people
The record shows who labeled and reviewed each file, which companies and tools were involved, and how each meets your security and ethical requirements. That covers our team, yours, and any third-party labeling tool the work passed through.
- 04
Deliver versions
Each delivery is a version you can diff, roll back, and cite, with its Data ID Card attached. If a file is corrected or removed later, the next version shows it, and earlier versions stay available to compare against.
Case study · Maritime technology company
Ocean video with rights and lineage on every file
A maritime technology company needed video of U.S. waters from vessels, shorelines, and elevated coastal positions, across sea states, weather, and light. The footage was sensitive: locations, vessels, and people had to stay unidentifiable, and the company needed to know it could use every clip.
Experienced boat operators and coastal observers from MLtwist's network recorded to a written protocol, and each video was checked and categorized by weather, water movement, visibility, and activity. MLtwist generated a Data Card for every video, recording capture location, data ownership, associated contracts, usage rights, and lineage.
The company received an organized, validated dataset tagged by environmental attributes and ready for AI development — with the transparency and traceability to use each clip in line with its terms. Early validation also meant fewer reshoots of ocean conditions that are hard to reproduce.
Read the case studyWhat you get
Answers for whoever asks where your data came from
Regulators, procurement reviewers, customers, and your own legal team.
-
The Data ID Card
Origin, transformations, people and vendors, and compliance, recorded for every file. It travels with the data, not in a separate document.
-
Versioned deliveries
Each delivery is a version you can diff, roll back, and cite. You can always say exactly which data a given model was trained on.
-
Your data stays yours
Files stay in your GCS, S3, or Azure bucket, encrypted at rest, with role-based access. The record describes your data without moving it somewhere else.
-
Built for public sector
Chain of custody for regulated programs, and purchasing through Carahsoft and Google Cloud Marketplace. The same record serves agencies, national labs, and commercial teams.
FAQ
Common questions
What's on a Data ID Card?
Where each file came from and under what terms, every step that changed it, who labeled and reviewed it, which companies and tools touched it, and how each met your requirements. For pre-labeled files, it also shows which model and prompt produced the draft.
Does it cover data we already have?
Yes. Data you bring gets a Data ID Card from the moment it enters MLtwist. Data we collect or generate gets one from capture.
Is it filled in by hand?
No. The record is written by the platform as the work runs — each Twist, review, and delivery adds to it. That's what makes it trustworthy when someone asks for it later.
Has it been used in published work?
Yes. Stanford NLP Group researchers used the Data ID Card to audit where their data came from and how it moved for two published studies. For published research, that record is part of the method.
Does it work for regulated programs?
Yes. Programs such as checkpoint-screening work for national labs need every step versioned and traceable to meet strict security requirements. The Data ID Card and versioned deliveries provide that chain of custody.
Who can see our data?
Access is role-based within your organization, and files stay encrypted at rest in your own storage.
Talk to us about data provenance & compliance
Tell us the data type, volume, and timeline. We'll scope it with our team, yours, or both — and deliver it versioned, in your format.
Also available through Carahsoft and Google Cloud Marketplace.