Solutions · Data preparation & curation

Raw data, cleaned and ready to label

Ingest straight from your cloud storage, clean and convert it, merge the metadata that gives it context, and pre-label it with models — so people spend their time correcting, not starting from zero.

Hands on a laptop showing a grid of image and video thumbnails, with two external drives and a cable on the desk

Overview

Most of a data program happens before anyone draws a label

Raw data rarely arrives ready to label. It comes from different teams, devices, and vendors, in mixed formats, with duplicates, corrupt files, and naming that made sense to whoever created it. The details that give a file meaning — a sensor reading, a capture location, a prior label — usually sit in another system. Before anyone labels anything, someone has to find the files, clean them, convert them, and match them with their context. That work is slow, it’s easy to get wrong, and it’s rarely anyone’s full-time job.

Good preparation does three things. It removes what shouldn’t be labeled, so budget goes to files that will improve the model. It puts every file in a form your tools can open, at a size and format that fits the work. And it brings context along, so annotators and reviewers see the whole picture instead of guessing. Done well, it also repeats: the next batch goes through the same steps, the same way, without an engineer rewriting a script.

MLtwist runs preparation as a pipeline of Twists — versioned steps that read from your Google Cloud Storage, Amazon S3, or Azure bucket and write the results back. Files are cleaned, split, resized, or transcoded, companion files and metadata are attached, and foundation models draft labels so people correct instead of starting from zero. The prepared files are organized into datasets and batches and sent to the labeling tool you choose. At Sandia National Laboratories, that approach cut screening-data preparation from eight weeks to three.

Use data preparation when raw files are slowing your program down, when annotators keep asking for context they can’t see, or when every new batch means another round of one-off scripts. It pairs with labeling — our team, yours, or both — but it also stands on its own when what you need is a clean, organized, well-documented dataset.

The problem

Why this is hard to do well

01

Messy inputs

Mixed formats, duplicates, corrupt files, and inconsistent naming arrive before a single label is drawn. Every one of them either wastes an annotator's time or slips into the training set as noise. Sorting it out by hand means someone opens files one at a time.

02

Context lives elsewhere

Metadata, sensor readings, and companion files sit in other systems, so annotators label blind. A body type, an item placement, or a prior label can change what the right answer is. When that context isn't next to the file, people guess, and the guesses end up in your data.

03

Weeks lost up front

Preparation is often the slowest part of a data program, and it's rarely anyone's full-time job. Engineers write one-off scripts between other work, and each new batch starts from scratch. Sandia found more than 75 places where errors could enter its process before it automated it.

How it works

What happens before the first label

Each step runs as a Twist — a versioned pipeline step — so the next batch is prepared exactly the same way.

  1. 01

    Connect your storage

    MLtwist reads directly from your Google Cloud Storage, Amazon S3, or Azure Blob bucket, so files stay where they are and nobody passes around hard drives. Each file is registered in MLtwist as it comes in, so it can be tracked from that point on.

  2. 02

    Clean and convert

    Corrupt and duplicate files are dropped, long captures are split, images are resized, large imagery is cut into tiles, and video is transcoded to formats labeling tools can play. Each operation is a Twist with fixed settings, so the same input always produces the same output.

  3. 03

    Merge context

    Metadata and companion files such as sensor readings, data cards, and prior labels are attached to each file, so annotators see the full picture. For multimodal programs, that means a scan, its images, and its metadata travel together instead of living in three systems.

  4. 04

    Pre-label

    Foundation models draft the labels before anyone opens the file. People correct those drafts instead of starting from a blank frame, which moves effort from drawing to judging. Pre-labels are kept alongside the source file, so you can see what the model proposed and what a person changed.

  5. 05

    Curate and send

    Files are organized into datasets and batches, and only what's worth labeling goes to the labeling tool — ours or yours. Near-duplicates and unusable files stay out of the queue, so labeling budget goes to data that will actually improve the model.

Case study · Retail analytics platform

Hundreds of thousands of products, sorted into one taxonomy

A retail analytics platform aggregates sales and inventory data from thousands of independent stores. Its product database came from many sources, each with its own names, descriptions, and categories, so similar items couldn't be compared. Many listings lacked the detail needed to categorize them at all.

MLtwist used AI models to propose a category and subcategory for every product, then had trained reviewers check each one against manufacturer sites and retailer listings, correcting the hallucinations and misclassifications the models introduced. MLtwist also set up a consistent categorization framework, with multi-step review across the whole catalog.

The result is one taxonomy that new products drop into as they arrive. Clean product data gave the platform more accurate analytics and reporting for retailers and brands, and the combination of AI speed and human validation delivered results much faster than manual work alone — with a repeatable process for every product that comes next.

products in one taxonomy, each checked by a person
100K+
Read the case study

What you get

Data that's ready to label, or ready to train on

  • Clean, consistent files

    Deduplicated, converted, and split to the formats and sizes your labeling and training tools expect. No one on your side has to reformat anything before the next step.

  • Context attached

    Metadata and companion files linked to each file, so no one labels blind. The same context is available to reviewers when they check the work.

  • Draft labels

    Model pre-labels kept alongside each source file, ready for people to correct. You keep both the draft and the final answer, not just the result.

  • A repeatable pipeline

    The same Twists run on the next batch, and each run records its version, inputs, and settings. When a result looks off months later, you can see exactly how it was prepared.

FAQ

Common questions

Do you need a copy of our data?

No. MLtwist reads from and writes back to your GCS, S3, or Azure bucket. Files stay in your storage, encrypted at rest, and access is role-based within your organization.

Which formats can you handle?

Images, video, audio, text, sensor captures, and 3D security scans, among others. Conversions are Twists, so a new format means a new pipeline step, not a new platform. Sandia's screening data, for example, uses the DICOS security imaging standard.

Can you prepare data for a labeling tool we already use?

Yes. Prepared batches can go to MLtwist's labeling tool or to third-party tools such as Datasaur, Kili, and Vertex AI. MLtwist can also pull the finished labels back out and post-process them for you.

What does pre-labeling actually save?

It changes the job from drawing every label to correcting a draft. How much that saves depends on the data. On unstable drone footage for one retailer, tuned pre-labeling helped cut labeling from about 10 minutes per frame to 7.

How much time does preparation save overall?

It depends on how messy the inputs are. Sandia National Laboratories cut screening-data preparation from eight weeks to three, and because the pipeline was modular, it could be changed when new issues appeared instead of rebuilt.

Do we have to label everything we prepare?

No. Preparation often ends with curation — choosing which files are worth labeling. Some programs use prepared data directly, such as a cleaned and standardized catalog or a structured set of emission activities.

Talk to us about data preparation & curation

Tell us the data type, volume, and timeline. We'll scope it with our team, yours, or both — and deliver it versioned, in your format.

Also available through Carahsoft and Google Cloud Marketplace.