arXiv:2504.05311cs.DBcs.CL2025-04

用简单查询语言从网页提取结构化数据,支持动态内容与脏数据处理。

Dr Web: a modern, query-based web data retrieval engine

  • 基于查询语言的模块化设计,灵活提取网页结构化数据。
  • 可处理动态加载内容与格式混乱的数据,提升提取鲁棒性。
  • 开源工具,适合需要自动化网页数据采集的研究与开发者。

本文介绍数据检索网络引擎(Data Retrieval Web Engine,简称DR Web),一种利用简单查询语言从网页中提取结构化数据的灵活、模块化工具。文章讨论了开发过程中解决的工程挑战,包括动态内容处理和杂乱数据提取。此外,还介绍了将DR Web引擎开源发布的步骤,突显其开放源代码潜力。

原文摘要 · Abstract (English)

This article introduces the Data Retrieval Web Engine (also referred to as doctor web), a flexible and modular tool for extracting structured data from web pages using a simple query language. We discuss the engineering challenges addressed during its development, such as dynamic content handling and messy data extraction. Furthermore, we cover the steps for making the DR Web Engine public, highlighting its open source potential.

数据提取网页抓取开源工具

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。