A Study on Effective Internet Data Extraction through Layout Detection

Sun Bok-Keun, & Han Kwang-Rok. (2005). A Study on Effective Internet Data Extraction through Layout Detection. INTERNATIONAL JOURNAL OF CONTENTS, 1(2), 5-9.

복사

Abstract

Currently most Internet documents including data are made based on predefined templates, but templates are usually formed only for main data and are not helpful for information retrieval against indexes, advertisements, header data etc. Templates in such forms are not appropriate when Internet documents are used as data for information retrieval. In order to process Internet documents in various areas of information retrieval, it is necessary to detect additional information such as advertisements and page indexes. Thus this study proposes a method of detecting the layout of Web pages by identifying the characteristics and structure of block tags that affect the layout of Web pages and calculating distances between Web pages. This method is purposed to reduce the cost of Web document automatic processing and improve processing efficiency by providing information about the structure of Web pages using templates through applying the method to information retrieval such as data extraction.

keywords: Information Retrieval, Data Extraction, Layout, HTML, XML Technologies

바로가기메뉴

논문 상세

Vol.1 No.2

A Study on Effective Internet Data Extraction through Layout Detection

Abstract

INTERNATIONAL JOURNAL OF CONTENTS