首页    期刊浏览 2024年09月20日 星期五
登录注册

文章基本信息

  • 标题:Deep Web Data Extraction by Using Vision-Based Item and Data Extraction Algorithms
  • 本地全文:下载
  • 作者:B.Sailaja ; Ch.Kodanda Ramu ; Y.Ramesh Kumar
  • 期刊名称:International Journal of Computer Science and Information Technologies
  • 电子版ISSN:0975-9646
  • 出版年度:2014
  • 卷号:5
  • 期号:6
  • 页码:7677-7682
  • 出版社:TechScience Publications
  • 摘要:Deep Web contents are accessed by queries submitted to Web databases and the returned data records are enwrapped in dynamically generated Web pages (they will be called deep Web pages in this paper). Extracting structured data from deep Web pages is a challenging problem due to the underlying intricate structures of such pages. Until now, a large number of techniques have been proposed to address this problem, but all of them have inherent limitations because they are Web-page-programming-language dependent. As the popular two-dimensional media, the contents on Web pages are always displayed regularly for users to browse. This motivates us to seek a different way for deep Web data extraction to overcome the limitations of previous works by utilizing some interesting common visual features on the deep Web pages. In this, a novel vision-based approach that is Web-page programming- language-independent is proposed. This approach primarily utilizes the visual features on the deep Web pages to implement deep Web data extraction, including data record extraction and data item extraction. We also propose a new evaluation measure revision to capture the amount of human effort needed to produce perfect extraction. Our experiments on a large set of Web databases show that the proposed vision-based approach is highly effective for deep Web data extraction
  • 关键词:Deep web extraction; vision based approach; enwrapped;item extraction; record extraction
国家哲学社会科学文献中心版权所有