Industries · Genealogy
Historical records, made searchable
Handwritten and printed birth, marriage, and other official records in many languages, pre-labeled by AI, verified by people, and delivered as JSON for search.
Overview
Turning archives into answers people can search
Genealogy runs on records: birth and marriage certificates, registers, and other official documents created over centuries by clerks who never imagined a search box. Most of them exist as images, handwritten or printed, in the language and layout of their place and time. Until the details on each record — names, dates, places, relationships — are pulled out and put into a consistent structure, the record can’t be found by the person searching for an ancestor.
Doing that by hand doesn’t scale. Researchers transcribing and verifying documents one at a time make slow progress through archives that hold millions of records, and fatigue brings errors. Fully automated reading has the opposite problem: it’s fast, but it stumbles on faded ink, unusual handwriting, and languages it wasn’t built for, and a single misread surname hides a record for good. The work needs machine speed and human judgment in the same process.
MLtwist combines them. A multilingual AI model pre-labels each record, drafting the fields so people correct rather than transcribe. Automated cleaning standardizes names, dates, and places so the same thing is written the same way everywhere. Reviewers verify the labeled data before it’s final, and the structured result is delivered as JSON that loads straight into a search database. For one leading genealogy company, that process made millions of historical records searchable.
It fits archives where the value is locked in document images and your users need to search them. The record images stay in your own cloud storage while the work runs, and nothing reaches your database until people have checked it. The same approach — AI extraction with human verification — also runs on modern documents, such as claims files for InsurTech.
The problem
What makes genealogy data hard
Handwriting and languages
Records span centuries, scripts, and languages, often in faded handwriting. The same kind of certificate can look completely different from one region or decade to the next. A system that reads one collection well can fail on the next.
Millions of documents
Manual transcription can't reach the scale of a global archive. Researchers painstakingly transcribing and checking records one at a time will never finish, and the error rate climbs as people tire. The backlog grows faster than any team can clear it.
Search depends on structure
Names, dates, and places must be standardized before anyone can find an ancestor. A misspelled surname or an inconsistent date format hides a record from the one person looking for it. An image of a record isn't searchable until its details are extracted.
How it works
From a scanned record to a searchable entry
- 01
Pre-label across languages
An AI model trained to recognize text in multiple languages reads handwritten and printed records and drafts the fields. People start from that draft instead of transcribing every line, which is what makes millions of records a realistic goal.
- 02
Clean and standardize
Extracted names, dates, and places are corrected and put into consistent formats. Automated cleaning fixes inconsistencies across records, so the same person or place is written the same way wherever it appears.
- 03
Verify with people
Reviewers check and validate the labeled data before it's final. A rigorous quality process catches what the model misread, which matters most in faded or unusual handwriting where automated reading is weakest.
- 04
Deliver as JSON
Structured records are delivered as JSON, ready to load into your search database. Each record arrives in the same structure, so it slots into your index without custom handling.
Case study · Genealogy company
Millions of official records, turned into a searchable database
A leading genealogy company holds vast collections of birth, marriage, and other official records, handwritten and printed, in many languages. Transcribing and verifying them by hand was slow, error-prone, and couldn't reach the scale of the archive.
MLtwist added a multilingual AI model to pre-label the extracted data, automated its cleaning and standardization, and verified every labeled record before delivering it as JSON. The structured output plugged straight into the company's searchable database.
The result is an extensive, highly accurate, searchable database of historical records from around the world. Millions of people use it to trace their family histories, find lost relatives, and uncover their ancestors' stories.
- of historical records made searchable
- Millions
What you get
Records your users can search
-
Structured fields
Names, dates, places, and other details extracted from each record. Every record follows the same structure, whatever its original layout.
-
Consistent formats
Standardized values, so searches match across records. A date or a place name is written one way, not ten.
-
Verified data
Human QA on labeled records before delivery. Nothing reaches your database until it has been checked.
-
JSON delivery
Ready to load into your database. No reformatting on your side before records go live.
More work
Related programs
FAQ
Common questions
Can you read handwritten records?
Yes. Handwritten and printed records are pre-labeled by AI models and then verified by people. The human review matters most on faded or unusual handwriting, where automated reading is least reliable.
Which languages?
Records from many regions and languages. The pre-labeling model was trained to process text in multiple languages, so a multilingual archive doesn't need a separate process for each one.
How do you keep records consistent?
Extracted fields go through automated cleaning and standardization, then human QA. Names, dates, and places come out in consistent formats that your search can match.
How do we receive the data?
As JSON, structured to load into your search database. Each delivery is a version, so you can see what changed between batches.
Where are the record images stored?
In your own cloud storage — GCS, S3, or Azure — encrypted at rest. MLtwist reads from and writes back to your bucket instead of keeping its own copy.
Bring us your genealogy data
Tell us the data type, volume, and timeline. We'll scope it with our team, yours, or both — and deliver it versioned, in your format.
Also available through Carahsoft and Google Cloud Marketplace.