Character string extraction from newspaper headlines with a background design by recognizing a combination of connected components

Hiroaki Takebe*, Yutaka Katsuyama, Satoshi Naoi

*Corresponding author for this work

Research output: Contribution to journalConference articlepeer-review

2 Citations (Scopus)

Abstract

In this paper we propose a new method of extracting a character string from images with a background design. In Japanese newspaper headlines, it is common for character components to be placed independent of background components. In view of this, we represent a character string candidate as a consistent combination of connected components, and we calculate its character string resemblance value. In this case, a character string resemblance value of a combination of connected components depends upon its character recognition result and the area of the rectangular area occupied by it. We then extract the combination of connected components that has the maximum character string resemblance value. We applied this method to 142 headline images. The results show that the method accurately extracted a character string from various kinds of images with a background design and the method has a favorable processing speed.

Original languageEnglish
Pages (from-to)22-29
Number of pages8
JournalProceedings of SPIE - The International Society for Optical Engineering
Volume3651
DOIs
Publication statusPublished - 1999
Externally publishedYes
EventProceedings of the 1999 6th Annual Conference on Document Recognition and Retrieval VI - San Jose, CA, USA
Duration: 1999 Jan 271999 Jan 28

ASJC Scopus subject areas

  • Electronic, Optical and Magnetic Materials
  • Condensed Matter Physics
  • Computer Science Applications
  • Applied Mathematics
  • Electrical and Electronic Engineering

Fingerprint

Dive into the research topics of 'Character string extraction from newspaper headlines with a background design by recognizing a combination of connected components'. Together they form a unique fingerprint.

Cite this