Digitization Decisions: Comparing OCR Software for Librarian and Archivist Use
https://journal.code4lib.org/articles/16132
Leanne Olson and Veronica Berry
This paper is intended to help librarians and archivists who are involved in digitization work choose optical character recognition (OCR) software. The paper provides an introduction to OCR software for digitization projects, and shares the method we developed
for easily evaluating the effectiveness of OCR software on resources we are digitizing.
We tested three major OCR programs (Adobe Acrobat, ABBYY FineReader, Tesseract) for accuracy on three different digitized texts from our archives and special collections at the University of Western Ontario. Our test was divided into two parts: a word accuracy
test (to determine how searchable the final documents were), and a test with a screen reader (to determine how accessible the final documents were). We share our findings from the tests and make recommendations for OCR work on digitized documents from archives
and special collections.
From: Issue 52 of the Code4Lib Journal.
The new issue is available at:
https://journal.code4lib.org/issues/issues/issue52
בברכה – דודו
יד יערי