Industries · AdTech

Brand-safety data for every format

Profanity and unsafe content flagged at the frame and the second, across on-screen text, speech, music, and images, plus IAB-aligned classifications that keep ads out of the wrong places.

A reviewer seen from behind at a desk facing a wall of small monitors showing video thumbnails

Overview

Training data for brand safety and content moderation

Advertisers pay to appear next to content, and they pay again when that content turns out to be offensive. Brand safety is the work of keeping ads away from harmful material; suitability is the finer judgment of which content fits a particular brand. The IAB’s guidelines give advertisers and platforms a shared vocabulary for both, and AdTech companies use AI to apply that vocabulary to text, images, video, and audio at a scale no team of human reviewers could cover.

Those models are only as good as the data behind them. A clip can look clean and still carry profanity in a caption, in dialogue, or buried in a song lyric, so training data has to cover every modality in the same piece of content. It also has to say exactly when a problem appears, down to the frame and the second, because brands and platforms need to know where the problem is, not just that one exists. And it has to be accurate in both directions: flag too much and good inventory goes unsold; flag too little and a brand lands next to something it can’t defend.

MLtwist builds this data with AI and people working together. Models pre-label every clip across modalities, trained reviewers confirm, correct, and time-stamp each flag, and multi-stage QA checks the result before delivery. Labels can follow IAB-aligned categories, your own taxonomy, or both. The same workflow produces training data for new classifiers and evaluation sets for measuring the ones you already run — see model evaluation for how those test sets are built.

Bring MLtwist in when you’re launching or retraining a brand-safety or content moderation model, when manual moderation can’t keep pace with the content coming in, or when you need to show how well a classifier actually performs. One AdTech company uses MLtwist for moderation training data flagged to the frame; another uses MLtwist datasets to assess its models against IAB guidelines.

The problem

What makes adtech data hard

01

Every modality at once

Offensive content hides in visuals, on-screen captions, spoken dialogue, and background lyrics — often in the same clip. A model trained on one modality misses the rest, so the training data has to cover all of them together.

02

Timestamp precision

Knowing a video contains something unsafe isn't enough. Brands and platforms need to know exactly when it appears, down to the frame and the second, and in which modality, so every label has to carry that detail.

03

False flags cost money

Over-flagging blocks good inventory that could have earned revenue. Under-flagging puts a brand next to content it can't defend. Both errors show up on the bottom line, so accuracy has to hold in both directions.

How it works

How a brand-safety dataset gets made

AI models make the first pass. People make the calls that decide where an ad can run.

  1. 01

    Ingest any source

    Content comes in as MP4 files or direct URLs from video platforms such as TikTok and YouTube. It's pulled in, cleaned, and normalized automatically, so reviewers work from consistent files whatever the source.

  2. 02

    Pre-label every modality

    AI models make a first pass over each clip, flagging candidate issues in the visuals, on-screen text, speech, and background music. Reviewers start from those flags instead of an empty timeline.

  3. 03

    Review to the frame

    Trained reviewers confirm, correct, or reject each flag. Every confirmed issue is marked with its exact frame and second and the modality it came from, so your model learns where the problem is as well as what it is.

  4. 04

    Map to the taxonomy

    Labels map to IAB-aligned brand-safety and suitability categories, your own taxonomy, or both. A single taxonomy keeps labels comparable across reviewers and batches, and lines your model's output up with categories your customers already use.

  5. 05

    QA and deliver

    Multi-stage QA checks the labels before anything ships. The result is delivered as structured data your models can train on or be scored against, versioned so one delivery can be compared with the next.

Case study · AdTech company

Profanity flagged to the frame across video, captions, speech, and music

A leading AdTech company needed a content moderation system that could catch offensive content in short-form and online video — in the visuals, on-screen captions, dialogue, and song lyrics — and mark exactly when it appears, to the frame and the second.

The content came from TikTok and YouTube, as MP4 files and direct URLs. MLtwist pre-labeled it with AI models, then ran its human-in-the-loop workflow for cleaning, transformation, labeling, and QA, with reviewers pinning down the exact timing and nature of each offensive moment.

The company reached 98% accuracy detecting offensive content, worked 45% faster than traditional manual moderation, and spent 40% less than it had on legacy moderation solutions. With frame-level flags across every modality, it built a multimodal content classification system that lets brands advertise with confidence and helps its own customers protect their reputations.

accuracy detecting offensive content
98%
faster than manual moderation
45%
Read the case study

What you get

Flags your ad stack can act on

  • Time-stamped flags

    Each unsafe moment is marked with its frame, second, and modality. Your model learns when and where the problem is, not just that a clip has one.

  • IAB-aligned labels

    Categories follow the industry's brand-safety and suitability guidelines, with your own taxonomy alongside if you have one. Labels stay consistent across reviewers and batches.

  • Fewer false flags

    Model output is reviewed by people to cut both over- and under-flagging. That protects your inventory and your advertisers' brands at the same time.

  • Evaluation sets

    Structured datasets for scoring your own classifiers, not just training them. One AdTech company uses MLtwist datasets to assess its brand-safety models against IAB guidelines.

FAQ

Common questions

Which formats can you review?

Video, images, text, and audio, including MP4 files and direct links from video platforms such as TikTok and YouTube. Mixed-modality content is reviewed as a whole, so a caption and the dialogue under it are checked together.

Can you follow our taxonomy?

Yes. Labels can follow the IAB brand-safety and suitability guidelines, your own taxonomy, or both. The taxonomy is set up before labeling starts, so every reviewer applies it the same way.

How precise are the flags?

Down to the frame and the second, with the modality recorded: visuals, on-screen text, speech, or music. Timestamp marking at that level was part of the spec in MLtwist's content moderation work for one AdTech company.

Is it cheaper than manual moderation?

In one program, AI pre-labeling plus human review ran 45% faster than traditional manual moderation and cost 40% less than legacy moderation solutions, at 98% accuracy detecting offensive content.

Can you help us measure the model we already run?

Yes. MLtwist builds structured, labeled datasets that AdTech companies use to assess model performance, so false positives and false negatives show up before they cost inventory or reputation. See model evaluation for how test sets are built and versioned.

Who does the reviewing?

Our trained reviewers, your team working in MLtwist's labeling tool, or both. The QA process and the delivery format stay the same whichever you choose.

Bring us your adtech data

Tell us the data type, volume, and timeline. We'll scope it with our team, yours, or both — and deliver it versioned, in your format.

Also available through Carahsoft and Google Cloud Marketplace.