arXiv:2502.14625cs.CLcs.IR2025-02被引 1

构建首个俄语新闻列表页数据集,提升多记录网页信息抽取效果。

Multi-Record Web Page Information Extraction From News Websites

  • 针对列表页设计多阶段抽取方法,适配复杂结构
  • 构建1.3万页俄语新闻列表数据集,规模超现有资源
  • 支持可选与多值属性,贴近真实网页场景

本文聚焦于包含多个记录的网页信息抽取问题,该任务在海量网络数据时代日益重要。尽管神经网络方法已显著提升网页信息抽取质量,但多数研究和数据集仍集中于详细页面,导致多记录“列表页”相对缺乏关注,尽管其应用广泛且具有实际意义。为此,我们构建了一个大规模、公开可用的列表页专用数据集,这是首个面向俄语的此类数据集。该数据集包含13,120个新闻列表网页,无论在规模还是复杂性上均超过现有资源。数据涵盖多种类型属性,包括可选和多值属性,更真实地反映现实世界列表页特征。这些特性使本数据集成为研究多记录网页信息抽取的重要资源。此外,我们提出了自己的多阶段信息抽取方法,并探索了MarkupLM在应对多记录网页特有挑战时的多种策略。实验验证了所提方法的有效性。通过公开数据集,我们旨在推动多记录网页信息抽取领域的发展。

原文摘要 · Abstract (English)

In this paper, we focused on the problem of extracting information from web pages containing many records, a task of growing importance in the era of massive web data. Recently, the development of neural network methods has improved the quality of information extraction from web pages. Nevertheless, most of the research and datasets are aimed at studying detailed pages. This has left multi-record "list pages" relatively understudied, despite their widespread presence and practical significance. To address this gap, we created a large-scale, open-access dataset specifically designed for list pages. This is the first dataset for this task in the Russian language. Our dataset contains 13,120 web pages with news lists, significantly exceeding existing datasets in both scale and complexity. Our dataset contains attributes of various types, including optional and multi-valued, providing a realistic representation of real-world list pages. These features make our dataset a valuable resource for studying information extraction from pages containing many records. Furthermore, we proposed our own multi-stage information extraction methods. In this work, we explore and demonstrate several strategies for applying MarkupLM to the specific challenges of multi-record web pages. Our experiments validate the advantages of our methods. By releasing our dataset to the public, we aim to advance the field of information extraction from multi-record pages.

信息抽取列表页俄语数据集MarkupLM

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。