Industries · Retail

Vision and catalog data for retail

Dense, safety-first labels on unstable drone-delivery video, and hundreds of thousands of products sorted into one taxonomy with AI speed and human checks.

A barcode scanner on a cart in a warehouse aisle lined with shelves of boxes

Overview

Two very different retail data problems

Retail AI runs on two kinds of data that have almost nothing in common. One is visual and safety-critical: delivery drones, store cameras, and warehouse systems that have to recognize people, pets, and objects in places they’ve never seen. The other is structured and enormous: product catalogs with hundreds of thousands of items that have to be named and categorized consistently before any analytics built on them can be trusted. Both are hard to get right, and both fail quietly when they’re wrong.

Drone delivery shows the visual side at its hardest. Before a drone drops a package in a backyard, it has to know whether a child, a dog, or a glass table is underneath. The footage it learns from is unstable — wind pushes the drone around in every frame — so every box has to be adjusted frame by frame, and each frame can take minutes to label. Catalog data fails differently. Products arrive from many sources with inconsistent names and categories, and while LLMs can sort them quickly, they also invent categories and misread products. A catalog that’s mostly right still skews every report built on it.

MLtwist handles each with the same principle: AI for speed, people for correctness. For drone video, we choose the labeling tool that copes best with high-motion footage, tune pre-labeling to it, and design a taxonomy that puts safety-critical classes first, with QA focused on catching every person and pet. For catalogs, AI models propose a category for every product, and trained reviewers verify each one against manufacturer and retailer listings before it’s accepted.

The result is data you can build on — frame-level labels that hold up through motion blur, and one product taxonomy that new items drop into as they arrive. For more on video work, see video and computer vision.

The problem

What makes retail data hard

01

Safety-critical vision

Delivery drones must spot people, pets, and fragile objects in backyards they've never seen. Every yard has a different layout, light, and mix of objects, and a missed detection can mean an injury or damaged property rather than a wrong recommendation.

02

Slow frames

Wind-blown drone footage shifts side to side and up and down in every frame. Each frame took about ten minutes to label properly, and at that pace productivity and consistency both suffer across a large video set.

03

Messy catalogs

Products from many sources use different names, descriptions, and categories, so the same item can look like several. LLMs can sort them fast, but on their own they hallucinate and misclassify — errors that flow straight into sales and inventory reports.

How it works

Retail's two hardest data problems

Unstable drone video and sprawling product catalogs, each handled with AI speed and human checks.

  1. 01

    Pick the right tool

    High-motion drone video breaks many labeling tools. We evaluate the options, choose the one that handles this footage best, and customize AI pre-labeling for the blur and camera movement, so annotators start from a strong draft instead of a blank frame.

  2. 02

    Label safety first

    The taxonomy puts humans, pets, and fragile objects first, because those are the classes the drone can't afford to miss. Consistency checks keep objects tracked accurately across frames even as the drone shifts, and QA is tuned for high recall on those classes.

  3. 03

    Let AI draft the catalog

    For product data, AI models propose a category and subcategory for every item at scale. That turns an impossible manual job — hundreds of thousands of products — into a review job people can finish.

  4. 04

    Verify against real sources

    Trained reviewers check each item against manufacturer sites and retailer listings, and correct the hallucinations and misclassifications the models introduced. Multi-step review keeps the result consistent, and the same workflow handles new products as they arrive.

Case study · Retail company

Safety-first labels for delivery-drone video

A large retailer building drone delivery needed its drones to identify people, pets, and objects in residential backyards before dropping a package. The footage was unstable — wind pushed drones side to side and up and down — and each frame took about ten minutes to label properly.

MLtwist chose the labeling tool that handled high-motion video best, customized AI pre-labeling for the unstable footage, and designed a taxonomy that put safety-critical classes first. Annotators used video workflows built for dynamic scenes, with proprietary consistency checks across frames, and multi-stage QA combined automated validation with expert review to keep recall high on people and pets.

Labeling time dropped from about ten minutes per frame to seven, a 30% cut in time and cost, while accuracy held despite motion blur. The retailer got consistent, production-ready datasets it can keep using as it trains and iterates on its drone-delivery models.

minutes of labeling per frame
10 → 7
lower labeling time and cost
30%
Read the case study

What you get

Retail data you can build on

  • Frame-level video labels

    Dense, consistent annotations tracked across frames, even through motion blur. Delivered in the format your training code reads.

  • High recall on safety classes

    People, pets, and fragile objects are checked first in QA. Those are the misses that cost the most, so they get the most scrutiny.

  • One product taxonomy

    Every item sits in a consistent category and subcategory, verified against outside sources. Similar products from different brands and retailers finally compare cleanly.

  • A process for new items

    The same AI-plus-review workflow categorizes new products as they arrive. The taxonomy stays consistent instead of drifting as the catalog grows.

FAQ

Common questions

Can you label unstable drone footage?

Yes. For one retailer, tuned pre-labeling and video workflows built for motion-heavy scenes cut labeling from about 10 minutes per frame to 7. Accuracy held despite the blur and camera movement.

How do you make sure the model doesn't miss people or pets?

The taxonomy puts safety-critical classes first, and QA is tuned for high recall on them. Automated validation and expert review both check those classes before a dataset is delivered.

Can't an LLM categorize our catalog on its own?

It can draft categories quickly, but it hallucinates and misclassifies. MLtwist pairs model suggestions with trained reviewers who check each item against manufacturer and retailer sources, and correct what the model got wrong.

How large a catalog can you handle?

Hundreds of thousands of products have gone into one taxonomy for a retail analytics platform that aggregates sales and inventory data from thousands of independent stores.

Can you collect retail data we don't have yet?

Yes. MLtwist runs real-world data collection, and generates synthetic data for cases you can't capture. Both get the same review as data you bring.

Bring us your retail data

Tell us the data type, volume, and timeline. We'll scope it with our team, yours, or both — and deliver it versioned, in your format.

Also available through Carahsoft and Google Cloud Marketplace.