Digitization Decisions: Comparing OCR Software for Librarian and Archivist Use

This paper is intended to help librarians and archivists who are involved in digitization work choose optical character recognition (OCR) software. The paper provides an introduction to OCR software for digitization projects, and shares the method we developed for easily evaluating the effectiveness...

Full description

Bibliographic Details
Main Authors: Leanne Olson, Veronica Berry
Format: Article
Language:English
Published: Code4Lib 2021-09-01
Series:Code4Lib Journal
Online Access:https://journal.code4lib.org/articles/16132
Description
Summary:This paper is intended to help librarians and archivists who are involved in digitization work choose optical character recognition (OCR) software. The paper provides an introduction to OCR software for digitization projects, and shares the method we developed for easily evaluating the effectiveness of OCR software on resources we are digitizing. We tested three major OCR programs (Adobe Acrobat, ABBYY FineReader, Tesseract) for accuracy on three different digitized texts from our archives and special collections at the University of Western Ontario. Our test was divided into two parts: a word accuracy test (to determine how searchable the final documents were), and a test with a screen reader (to determine how accessible the final documents were). We share our findings from the tests and make recommendations for OCR work on digitized documents from archives and special collections.
ISSN:1940-5758