Historical newspapers dataset released

18 views
Skip to first unread message

Eben English

unread,
Sep 4, 2026, 10:58:11 AM (4 days ago) Sep 4
to ai4lam
In collaboration with Boston Public Library, Harvard Law Library's Institutional Data Initiative has released a data set on HuggingFace derived from the BPL's public domain newspapers collection.

https://huggingface.co/collections/institutional/institutional-newspapers

This data set includes 1,473,635 newspaper scans from issues published between 1795 and 1930, which have been processed using ML and AI techniques to enhance the OCR and segment each article into distinct semantic units. This output has produced 83,147,041 individual crops segmented from those scans, over 16 billion o200k_base tokens of VLM OCR text, as well as bounding box coordinates, raw OCR, text analysis, crop type classification, language detection, named-entity recognition, subject classification, reading order detection, and text + image vector embeddings for each segment.

The publication of this data set will be of significant use for computational linguistics, AI model training, and historical research. The data processing pipeline created through this project can also serve as a model for other organizations that want to enhance collections of historical newspapers with improved full-text searching and subject analysis.

The dataset, pipeline code, and technical documentation are all openly available:

Thanks,

Eben English
Boston Public Library
Reply all
Reply to author
Forward
0 new messages