Multi-Scale Multi-Task FCN for Semantic Page Segmentation and Table Detection

Dafang He, Scott Cohen, Brian Price, Daniel Kifer, C. Lee Giles

Research output: Chapter in Book/Report/Conference proceedingConference contribution

18 Scopus citations

Abstract

Page segmentation and table detection play an important role in understanding the structure of documents. We present a page segmentation algorithm that incorporates state-of-The-art deep learning methods for segmenting three types of document elements: text blocks, tables, and figures. We propose a multi-scale, multi-Task fully convolutional neural network (FCN) for the tasks of semantic page segmentation and element contour detection. The semantic segmentation network accurately predicts the probability at each pixel of the three element classes. The contour detection network accurately predicts instance level 'edges' around each element occurrence. We propose a conditional random field (CRF) that uses features output from the semantic segmentation and contour networks to improve upon the semantic segmentation network output. Given the semantic segmentation output, we also extract individual table instances from the page using some heuristic rules and a verification network to remove false positives. We show that although we only consider a page image as input, we produce comparable results with other methods that relies on PDF file information and heuristics and hand crafted features tailored to specific types of documents. Our approach learns the representative features for page segmentation from real and synthetic training data. %, and produces good results on real documents. The learning-based property makes it a more general method than existing methods in terms of document types and element appearances. For example, our method reliably detects sparsely lined tables which are hard for rule-based or heuristic methods.

Original languageEnglish (US)
Title of host publicationProceedings - 14th IAPR International Conference on Document Analysis and Recognition, ICDAR 2017
PublisherIEEE Computer Society
Pages254-261
Number of pages8
ISBN (Electronic)9781538635865
DOIs
StatePublished - Jul 2 2017
Event14th IAPR International Conference on Document Analysis and Recognition, ICDAR 2017 - Kyoto, Japan
Duration: Nov 9 2017Nov 15 2017

Publication series

NameProceedings of the International Conference on Document Analysis and Recognition, ICDAR
Volume1
ISSN (Print)1520-5363

Other

Other14th IAPR International Conference on Document Analysis and Recognition, ICDAR 2017
CountryJapan
CityKyoto
Period11/9/1711/15/17

All Science Journal Classification (ASJC) codes

  • Computer Vision and Pattern Recognition

Fingerprint Dive into the research topics of 'Multi-Scale Multi-Task FCN for Semantic Page Segmentation and Table Detection'. Together they form a unique fingerprint.

  • Cite this

    He, D., Cohen, S., Price, B., Kifer, D., & Giles, C. L. (2017). Multi-Scale Multi-Task FCN for Semantic Page Segmentation and Table Detection. In Proceedings - 14th IAPR International Conference on Document Analysis and Recognition, ICDAR 2017 (pp. 254-261). (Proceedings of the International Conference on Document Analysis and Recognition, ICDAR; Vol. 1). IEEE Computer Society. https://doi.org/10.1109/ICDAR.2017.50