Solutions · AI workloads & pipelines

Run AI workloads on your data, where it lives

Twists are containerized pipelines that run on your data where it lives: model pre-labeling, LLM-powered ETL, format conversion, and exports to the tools you already use. Use ours, or build your own in the browser with the Twist AI Builder.

A laptop on a steel cart in an empty server aisle, its screen showing a pipeline diagram of connected boxes

Overview

Turn one-off data scripts into pipelines your whole team can run

Most AI data work is a chain of small jobs: pull files from storage, convert them, run a model over them, reshape the output, push it to a labeling tool, pull the labels back, and package the result. In many teams each of those jobs is a script someone wrote for one dataset. The scripts work until the format changes, the volume grows, or the person who wrote them moves on — and nobody can say afterward which version of which script touched which file.

A pipeline you can rely on looks different. Each step is packaged so it runs the same way every time, on any batch, for anyone on the team. Its inputs and settings are declared, not buried in code. It runs where the data lives instead of on a laptop. And every run leaves a record of the version, the person who started it, and the files it touched, so a result can be traced and reproduced months later.

In MLtwist those steps are called Twists. A Twist is a container with declared inputs and outputs, published to your organization as a version. Use ready-made Twists for import, conversion, pre-labeling, and export, or build your own in the Twist AI Builder, a browser editor with an AI assistant and test builds. Twists read from and write back to your GCS, S3, or Azure storage, use stored secrets instead of pasted credentials, and connect to labeling tools such as Datasaur, Kili, and Vertex AI. Language models can run inside them too — see using LLMs for ETL on AI data.

Use Twists when your data work has outgrown notebooks, when the same processing has to run on every new batch, or when you need to show how a training set was produced. They’re the engine behind data preparation and labeling at MLtwist, and you can run them on your own data with your own team.

The problem

Why this is hard to do well

01

Glue code everywhere

Every dataset gets its own scripts, and they live on one engineer's laptop. When that person is out, or the format changes, the pipeline stops. Nobody else knows which version of the script produced last quarter's training set.

02

Scale

Hundreds of thousands of files, GPU models, and long-running jobs don't fit in a notebook. Work that ran fine on a sample stalls on the full dataset. One program ran nine processing stages and 100,000 file transformations — far past what a folder of scripts can manage.

03

Traceability

Months later, nobody can say which code, version, or settings touched which file. When a model regresses, you can't tell whether the data changed or the processing did. Audits and customers ask the same question, and the answer is a guess.

How it works

How a Twist runs

A Twist is a container with declared inputs and outputs. Once it's published, anyone on your team can run it on a batch.

  1. 01

    Pick or build a Twist

    Start from ready-made Twists for import, conversion, pre-labeling, and export. When your data needs something they don't do, write your own in the Twist AI Builder, a browser editor where an AI assistant helps you write and edit the code.

  2. 02

    Test and publish

    Test-build the container before anyone relies on it. Publishing builds it and adds it to your organization as a versioned Twist that everyone can see. Earlier versions stay available, so a change never silently replaces the step last month's data went through.

  3. 03

    Configure the run

    Each Twist declares its settings, and the run form is built from them. Fill them in from shared variables and stored secrets instead of pasting credentials, so the same Twist can point at different buckets or models without anyone copying keys around.

  4. 04

    Run it on a batch

    Send a dataset or batch to the Twist, choosing the Twists assigned to your project. It runs as a containerized job, reading from and writing to your cloud storage, so the data never has to pass through someone's laptop.

  5. 05

    Trace every run

    The run history records the Twist version, who started it, how many files it touched, and its settings. When a result looks wrong, you can find the exact run that produced it and rerun or fix that step.

Case study · Defense company

Raw radar tiles in, the customer's exact format out

A defense company training vessel-detection models needed thousands of synthetic aperture radar (SAR) images labeled with tilted bounding boxes — sometimes more than 100 ships in one image. The raw data arrived as large tiles, and the labels had to come back in a specific, non-generic format.

MLtwist built the pipeline around the labeling. Tiles were split and resized into consistent, annotation-ready images automatically, AI-assisted tools placed angled boxes along each ship's heading, and the output was written straight into the customer's format, with every annotation version-controlled.

Annotation time dropped by 30%, and project timelines went from months to weeks despite the density of the images. Because the pipeline did the preprocessing and packaging, the data went into model training with no pre- or post-processing on the customer's side, and the tilted boxes gave the models a more realistic fit to each vessel.

less annotation time
30%
post-processing steps before training
0
Read the case study

What you get

Pipelines you own, not scripts on a laptop

  • Versioned Twists

    Each published Twist is a versioned container your whole organization can run. New versions don't overwrite old ones, so past results stay reproducible.

  • Custom steps, built with you

    Twists written in the Twist AI Builder and test-built before anything is published. Our team can build them for you, or your engineers can build them themselves.

  • Connections to your stack

    GCS, S3, and Azure storage, labeling tools such as Datasaur, Kili, and Vertex AI, and a REST API for uploads and downloads. Data moves between them without manual exports.

  • A run history

    Who ran what, which version, on how many files, with which settings — kept for every run. It's the record you need when a result has to be explained.

FAQ

Common questions

What is a Twist?

A containerized pipeline step that runs on your data: import, clean, convert, pre-label, send to a labeling tool, pull labels back, or deliver. Twists chain together into a full pipeline, and each one is versioned when it's published.

Can we write our own?

Yes. Describe what the step should do and let the AI assistant draft it, or write the code yourself. Test-build it, then publish it as a versioned Twist that anyone in your organization can run.

Do we need engineers to run a Twist?

No. Once a Twist is published, running it means choosing a batch and filling in its settings. Engineering is only needed to build a new step, and the Twist AI Builder shortens that too.

Can a pipeline use LLMs?

Yes. LLM-powered ETL Twists reshape messy text and metadata into the schema your training code expects. Model pre-labeling Twists draft labels on images, video, and text before people review them.

How are credentials handled?

Stored as secrets and selected when a Twist runs, so no one pastes keys into settings. Data is read from and written back to your own storage.

How is this different from data preparation?

Twists are the mechanism; preparation is one job they do. The same Twists also run exports, deliveries, format conversions, and custom model steps — anything you'd otherwise script by hand.

Talk to us about ai workloads & pipelines

Tell us the data type, volume, and timeline. We'll scope it with our team, yours, or both — and deliver it versioned, in your format.

Also available through Carahsoft and Google Cloud Marketplace.