Table Extraction from Web Pages Using Conditional

Abstract

Table is one of the ways to visualize information on web pages. The abundant number of web pages that compose the World Wide Web has been the motivation of information extraction and information retrieval research, including the research for table extraction. Besides, there is a need for a system which is designed to specifically handle location-related information. Based on this background, this research is conducted to provide a way to extract location-related data from web tables so that it can be used in the development of Geographic Information Retrieval (GIR) system. The location-related data will be identified by the toponym (location name). In this research, a rule-based approach with gazetteer is used to recognize toponym from web table. Meanwhile, to extract data from a table, a combination of rule-based approach and statistical-based approach is used. On the statistical-based approach, Conditional Random Fields (CRF) model is used to understand the schema of the table. The result of table extraction is presented on JSON format. If a web table contains toponym, a field will be added on the JSON document to store the toponym values. This field can be used to index the table data in accordance to the toponym, which then can be used in the development of GIR system.

Details

Title

Table Extraction from Web Pages Using Conditional Random Fields to Extract Toponym Related Data

Author

Hayyu’ Luthfi Hanifah¹; Akbar, Saiful¹

¹ School of Electrical Engineering and Informatics, Bandung Institute of Technology

Publication year

2017

Publication date

Jan 2017

Publisher

IOP Publishing

ISSN

17426588

e-ISSN

17426596

Source type

Scholarly Journal

Language of publication

English

DOI

https://doi.org/10.1088/1742-6596/801/1/012064

ProQuest document ID

2573811055

© 2017. This work is published under http://creativecommons.org/licenses/by/3.0/ (the “License”). Notwithstanding the ProQuest Terms and Conditions, you may use this content in accordance with the terms of the License.

Table Extraction from Web Pages Using Conditional Random Fields to Extract Toponym Related Data

Jump to:

Abstract

Details

Suggested sources