Industries · Research
Research-grade data for language and science
Datasets for published research, from the first Universal Dependencies dataset for Sindhi to a study of NER on Global Englishes, with the record of sources and people a methods section needs.
Overview
Datasets that hold up in a methods section
Research datasets carry a burden commercial training data doesn’t: they have to survive peer review, and often they become the benchmark other people build on. That changes what “good” means. The labels have to follow a formal framework consistently. The annotators have to be qualified for the language or domain, not just available. And the paper has to say exactly where the data came from, who labeled it, which tools were used, and how quality was checked. A methods section can’t cite a process no one wrote down.
The practical work around a research dataset is where time goes. Someone has to find native speakers for a language with few speakers online, configure a labeling tool for a framework like Universal Dependencies, run revision rounds until agreement is acceptable, and turn tool exports into something releasable. Often that someone is a graduate student, and none of it advances the research question.
MLtwist takes on that work for university research groups and national labs. We recruit annotators from our partner network, set up and run projects in tools such as Datasaur and Kili, drive revision rounds with automated quality checks, and post-process finished labels into the form the researchers need. Every dataset gets a Data ID Card recording its sources, tools, companies, and people, so provenance is part of the method rather than an afterthought. The researchers design the study and own the results.
That’s how Stanford NLP Group researchers built the first Universal Dependencies dataset for Sindhi, with two native-speaker annotators who became co-authors, and ran a study showing that English named entity recognizers struggle on English written by non-native speakers. Beyond language, MLtwist won a U.S. Department of Energy award to build PRIMED, a platform for organizing and labeling microscopy data, and partnered with Lawrence Berkeley National Laboratory on microscopy data tools.
The problem
What makes research data hard
Annotators who hold up to review
Low-resource language work needs native speakers who can annotate to a formal linguistic framework, not just read the text. Those people are hard to find, and their work has to survive peer review. A general annotation crowd rarely has either the language or the rigor.
A methods section to fill
Reviewers ask where the data came from, who labeled it, which tools were used, and how quality was checked. Reconstructing that after the fact is slow and incomplete. The record has to be kept while the work happens, or it won't exist when the paper is written.
Tool work eats research time
Configuring labeling tools, preparing data for them, running revision rounds, and post-processing the labels all take time. None of it is research, and it often falls on graduate students. Every week spent on tooling is a week not spent on the question.
How it works
How a research dataset gets made
The researchers design the study and own the results. MLtwist handles the people, the tools, and the record.
- 01
Recruit annotators
Native speakers and domain experts come from MLtwist's partner network, chosen for the language and the annotation framework. For a low-resource language, finding qualified speakers is often the hardest part of the project, so it starts first and runs alongside study design.
- 02
Set up the tools
MLtwist prepares the data and configures projects in labeling tools such as Datasaur and Kili, so the team labels consistently from the first sentence. Researchers agree on the guidelines; MLtwist handles the setup, imports, and access, so no one on the research team has to become a tool administrator.
- 03
Run revision rounds
Automated quality checks drive each human-in-the-loop revision round, flagging items that break the guidelines or disagree across annotators. Annotators fix what the checks catch, and the loop repeats until the labels meet the bar the researchers set.
- 04
Post-process
Finished labels are pulled from the labeling tool and processed into the form the researchers need, automatically. For a treebank, that means output that fits the framework's format and can be released as a dataset, not a tool export that someone has to clean by hand.
- 05
Document it
A Data ID Card records the sources, tools, companies, and people behind every dataset, along with how each met ethical and security requirements. Researchers use it to audit how the data moved, and it gives the methods section a record to cite.
Case study · Stanford
The first Universal Dependencies dataset for Sindhi
Sindhi is spoken by about 40 million people in India and Pakistan but has few labeled datasets or pretrained embeddings. Stanford NLP Group researchers set out to build its first modern NLP foundation using the Universal Dependencies framework, the standard for annotating grammar consistently across languages.
MLtwist recruited two native Sindhi speakers from its partner network — both co-authors on the paper — prepared the data and set up the labeling projects in Datasaur and Kili, ran automated QA through the human-in-the-loop revision rounds, and post-processed the finished labels for the researchers.
The study analyzed about 6,000 Sindhi sentences and released a dataset in Universal Dependencies versions 2.16 and 2.17. The researchers used the Data ID Card to audit where the data came from and how it moved, which companies and tools transformed it, and how each met ethical and security requirements.
- Sindhi sentences in the released dataset
- ~6,000
- published Stanford NLP studies supported
- 2
What you get
A dataset you can publish
-
Framework-ready annotations
Labels in the form your framework needs, such as Universal Dependencies. Ready to analyze and release, not a raw export from a labeling tool.
-
A documented process
Guidelines, revision rounds, and quality checks recorded as the work happens. Everything a reviewer asks about in the methods section is already written down.
-
A Data ID Card
Sources, tools, companies, and people recorded for every dataset. It shows how the data moved and how each party met ethical and security requirements.
-
Tooling handled
Projects set up and run in the labeling tool that suits the task, such as Datasaur or Kili. Your team works on the research, not on tool configuration.
More work
Related programs
Retail analytics platform · Retail
How MLtwist Supported a Retail Analytics Platform in Structuring Product Data at Scale
Cleantech company · CleanTech
How MLtwist Supported a Cleantech Company Tracking Carbon Emission Activity for Regulatory Action
FAQ
Common questions
Do you work with universities?
Yes. Stanford's Institute for Human-Centered Artificial Intelligence (HAI) selected MLtwist as a technology provider in 2023. Since then MLtwist has supported two published Stanford NLP Group studies: the Sindhi Universal Dependencies dataset and a study of named entity recognition on Global Englishes.
Can the annotators be credited?
Yes. On the Sindhi study, both native-speaker annotators MLtwist recruited are co-authors on the paper. How contributors are credited is the research team's call; MLtwist records who did what so it's easy to get right.
What did the Global Englishes study find?
It asked whether named entity recognizers trained on native-speaker English work as well on English written by proficient non-native speakers. They don't — differences as basic as how names are structured hurt performance. MLtwist preprocessed the data, set it up in Datasaur, ran automated QA through the revision rounds, and post-processed the results.
Do you work with national labs?
Yes. MLtwist won a U.S. Department of Energy award to build PRIMED, a platform for organizing and labeling microscopy data, and partnered with Lawrence Berkeley National Laboratory on tools that help researchers use fast-growing microscopy data.
Who owns the data and the results?
You do. MLtwist works on your data, with your guidelines, and delivers the labeled dataset back to you. The Data ID Card records the terms each file came with, so release decisions are straightforward.
Can you work in the labeling tool our lab already uses?
Often, yes. MLtwist has set up and run research projects in Datasaur and Kili, and can run work in other third-party tools, then pull, check, and post-process the results.
Bring us your research data
Tell us the data type, volume, and timeline. We'll scope it with our team, yours, or both — and deliver it versioned, in your format.
Also available through Carahsoft and Google Cloud Marketplace.