构建实时网页信息抽取基准,评估系统在动态网站上的真实表现。
LiveWeb-IE: A Benchmark For Online Web Information Extraction
- 设计动态基准,用实时网站和多层级查询测试抽取能力。
- 提出视觉定位爬虫框架,可精准定位页面内容并提取图文链接。
- 适合研究网页抽取、智能代理或需要高鲁棒性的应用开发者。
网页信息抽取(WIE)旨在自动从网页中提取数据,具有广泛应用价值。传统评估依赖单时间点的网页快照,无法反映网页随时间变化的特性,导致模型性能难以泛化到真实场景。为此,我们提出 dataset,一个面向实时网站的新型基准,基于受信任且授权的网站,设计涵盖文本、图像、超链接等多类数据的自然语言查询,并按属性数量与基数划分为四个复杂度等级,实现对WIE系统的细粒度评估。同时,我们提出视觉定位爬虫(VGS),一种模拟人类认知过程的多阶段智能体框架,通过视觉聚焦缩小目标区域,高效提取所需信息。在多种主干模型上的大量实验表明,VGS具备优异的性能与鲁棒性。本研究为构建实用、可靠的WIE系统奠定了基础。
原文摘要 · Abstract (English)
Web information extraction (WIE) is the task of automatically extracting data from web pages, offering high utility for various applications. The evaluation of WIE systems has traditionally relied on benchmarks built from HTML snapshots captured at a single point in time. However, this offline evaluation paradigm fails to account for the temporally evolving nature of the web; consequently, performance on these static benchmarks often fails to generalize to dynamic real-world scenarios. To bridge this gap, we introduce \dataset, a new benchmark designed for evaluating WIE systems directly against live websites. Based on trusted and permission-granted websites, we curate natural language queries that require information extraction of various data categories, such as text, images, and hyperlinks. We further design these queries to represent four levels of complexity, based on the number and cardinality of attributes to be extracted, enabling a granular assessment of WIE systems. In addition, we propose Visual Grounding Scraper (VGS), a novel multi-stage agentic framework that mimics human cognitive processes by visually narrowing down web page content to extract desired information. Extensive experiments across diverse backbone models demonstrate the effectiveness and robustness of VGS. We believe that this study lays the foundation for developing practical and robust WIE systems.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。