Training data equity

Training Data Equity: Who Gets Into the Corpus

Training data equity is the principle that the corpora used to train AI systems should represent the communities those systems will describe and serve. It is decided by digitization budgets and archival priorities rather than by model design, which means the composition of tomorrow's models is being set today by whoever chooses which paper records to scan.

The archive is the model

What a language model can say about a neighborhood, a newspaper, or a religious tradition is fixed long before training, by whether that material entered machine-readable form. Where it did not, the model reconstructs from outside accounts written by people who were not there.

This is the operating premise of the Living Archive Series: recovering degraded microfilm and unscanned runs is not nostalgia work, it is corpus work. Every page restored is a page a future model can read.

A pipeline that does not launder errors

Capture at a standard high enough to re-run later. Machine transcription for speed. Model-assisted correction constrained by a domain glossary of names, places, and vernacular so the corrector cannot normalize them away. Quality sampling against the original page. A final human reader who knows the community. Removing the last step is what turns a restoration into a fabrication.

Frequently asked

Why does training data equity matter more than model tuning?
Tuning adjusts how a model behaves with what it knows. Equity in the training corpus determines what it can know at all. Missing sources cannot be tuned back into existence.
How do you improve training data equity in practice?
Digitize under-represented archives to a re-runnable standard, correct OCR with a domain glossary rather than generic spellcheck, sample against original pages, and keep a human reader from the relevant community at the final gate.

Read the books behind this

Published titles by Robert Shumake that develop this material at length.