Industries · Autonomy
Real-world and synthetic data for autonomous systems
Cars, utility trucks, and vessels all need data from conditions they rarely meet in testing. MLtwist collects it on the road and at sea, generates the rest synthetically, and tags every clip by condition.
Overview
Why autonomy data is a coverage problem
Autonomous systems fail at the edges of what they’ve seen. A car that drives well in daylight can struggle in fog. A vessel’s perception model trained on calm water can misread glare, rough seas, or an unfamiliar coastline. The conditions that cause trouble — night, rain, a near miss, a crowded harbor — are exactly the ones a test fleet meets least. That makes autonomy data a coverage problem before it’s a labeling problem: the question is less “how many hours of footage” and more “which conditions are missing.”
Two things make coverage hard to close. The first is that real data is slow and uneven. Fleets drive the routes they drive, and waiting for rare conditions to happen on their own can take longer than the program has. The second is that autonomy data has to behave correctly, not just look correct. Planning and navigation models learn from how vehicles move — how often a utility truck stops, how it reverses to a curb, how timing changes at rush hour — so generated data with the wrong motion does real damage.
MLtwist does both, and plans them together. For conditions that can be captured safely, vetted drivers and boat operators record to a written spec with cameras positioned to match your sensors, and every upload is checked. For conditions that can’t, a small set of real routes is expanded synthetically with controlled light and weather, and automated checks validate stop frequency, direction, and timing. Every clip is tagged by condition, so coverage is something you can measure, and every file carries a Data ID Card.
This works for passenger cars, utility trucks, and maritime systems alike. For how the collection side runs, see data collection; for the generation side, see synthetic data generation.
The problem
What makes autonomy data hard
The long tail
Fog, glare, night, rough seas, and near misses are rare in fleet data, but that's where perception models fail. A fleet can drive for months without meeting the one condition that matters most — and a model that has never seen it has no way to handle it.
Behavior, not just pixels
Autonomy data has to move right, not just look right. Stop frequency, reverse maneuvers, curbside approaches, turning angles, and timing all shape what a planning model learns — and a route that looks real but behaves oddly teaches it the wrong habits.
Sensitive traces
Real routes reveal where people live, work, and travel, and ocean footage can identify vessels and the people aboard them. Locations, vessels, and people must stay unidentifiable, which limits how much real data you can keep and share.
How it works
Real data where you can get it, synthetic where you can't
Most autonomy programs need both. The work is deciding which conditions to capture and which to generate, then making the two sets agree.
- 01
Plan coverage
Start from the scenarios your model must handle — regions, road or water types, times of day, weather, and events — and decide which to collect and which to generate. Common, safe conditions are usually cheaper to record; rare and dangerous ones are generated.
- 02
Collect to sensor spec
Vetted drivers and boat operators record with cameras placed to match your vehicle's sensors, whether that's a windshield mount at a fixed height or a deck-level view of the horizon. Every upload is checked before it counts.
- 03
Expand synthetically
A few real routes become a broad synthetic set with controlled light and weather. Automated checks validate stop frequency, direction, speed patterns, and timing, so the generated routes behave like the real ones they were built from.
- 04
Tag by condition
Every clip is categorized by conditions such as weather, visibility, water movement, and activity. Coverage stops being a guess — you can see which conditions are thin before a model is trained on the set.
- 05
Deliver versions
Each delivery is a version you can diff and roll back, with a Data ID Card for every file. When a model's behavior changes, you can trace which data it was trained on.
Case study · Autonomous trucking company
Neighborhood routes for autonomous utility trucks, from a few real samples
An autonomous trucking company needed training data for utility trucks in residential neighborhoods — frequent stops, reverse maneuvers, curbside approaches, and road layouts, driveway spacing, and traffic flow that change from one neighborhood to the next. Real data covered only a handful of routes, and it couldn't expose real locations.
MLtwist built a route-generation pipeline that turned those samples into a broad synthetic set with realistic movement and stopping behavior. Morning light, midday glare, fog, and rain were layered in through controlled conditioning, and automated checks validated stop frequency, direction, and timing. New neighborhood types were added as the models evolved.
The result is a scalable synthetic dataset that reflects how utility trucks actually move through American neighborhoods. The company's autonomy team can test perception, planning, and navigation models under safe, reproducible conditions, with fewer data bottlenecks as the models change.
Read the case studyWhat you get
Coverage you can show
-
Condition-tagged video
Every clip is tagged by time of day, weather, and scene. Gaps in coverage are visible before training, not after a field failure.
-
One consistent viewpoint
Real footage is recorded from the camera position your sensors use and checked on every upload. The model learns the road or the water, not the differences between cameras.
-
Synthetic expansion
Generated routes and scenes keep your real data's viewpoint and behavior. They fill in the rare and risky conditions without drifting from what your vehicles actually see.
-
Frame-level labels
When the program needs annotations, objects are labeled and tracked across frames and reviewed as video. Drift and ID switches get caught before delivery.
More work
Related programs
Maritime technology company · Autonomy
Maritime Company Uses MLtwist for Nationwide Video Data Collection
Automotive company · Autonomy
Automotive Company Uses MLtwist for Nationwide Commute Video Data Collection
Technology company · Autonomy
MLtwist Generates Synthetic Commute Data for Autonomous Driving
FAQ
Common questions
Should we collect or generate?
Usually both. Collect the common conditions and the viewpoints your sensors see; generate the rare, dangerous, and private cases. Real samples also anchor the synthetic set, so the generated data stays true to how your vehicles actually operate.
Which kinds of autonomy do you work on?
Past programs cover passenger-car commutes, autonomous utility trucks in residential neighborhoods, and maritime systems. MLtwist has also labeled drone video for a retailer's delivery drones and for defense programs.
Do you work on maritime autonomy?
Yes. MLtwist collected ocean video across U.S. waters from vessels, shorelines, and elevated coastal positions, covering sea states, weather, light, and water colors. Each video was categorized by weather, water movement, visibility, and activity.
How do you check that synthetic routes behave realistically?
Automated validation checks behavior, not just appearance — stop frequency, directionality, speed patterns, timestamps, and route consistency. Routes that look right but move wrong are caught before delivery.
How do you protect location privacy?
Collected footage is anonymized where the program requires it, and synthetic routes let models train without identifiable location data. Each file's Data ID Card records ownership and usage rights.
Bring us your autonomy data
Tell us the data type, volume, and timeline. We'll scope it with our team, yours, or both — and deliver it versioned, in your format.
Also available through Carahsoft and Google Cloud Marketplace.