Context Dependent Thresholding and Filter Selection for Optical Character Recognition

Thresholding algorithms and filters are of great importance when utilizing OCR to extract information from text documents such as invoices. Invoice documents vary greatly and since the performance of image processing methods when applied to those documents will vary accordingly, selecting appropriat...

Full description

Bibliographic Details
Main Author: Kieri, Andreas
Format: Others
Language:English
Published: Uppsala universitet, Institutionen för informationsteknologi 2012
Subjects:
Online Access:http://urn.kb.se/resolve?urn=urn:nbn:se:uu:diva-197460
Description
Summary:Thresholding algorithms and filters are of great importance when utilizing OCR to extract information from text documents such as invoices. Invoice documents vary greatly and since the performance of image processing methods when applied to those documents will vary accordingly, selecting appropriate methods is critical if a high recognition rate is to be obtained. This paper aims to determine if a document recognition system that automatically selects optimal processing methods, based on the characteristics of input images, will yield a higher recognition rate than what can be achieved by a manual choice. Such a recognition system, including a learning framework for selecting optimal thresholding algorithms and filters, was developed and evaluated. It was established that an automatic selection will ensure a high recognition rate when applied to a set of arbitrary invoice images by successfully adapting and avoiding the methods that yield poor recognition rates.