仅用自有的多次抓取数据,估算网页存档的绝对覆盖率。
Estimating Absolute Web Crawl Coverage From Longitudinal Set Intersections
- 通过分析连续抓取间的网页重叠,构建简单的瓮模型来估计覆盖率。
- 在德国学术网络15次半年度抓取中,稳定阶段覆盖率达46%。
- 无需外部真实数据,适用于任何纵向聚焦抓取场景。
网络存档保存了部分网页,但量化其完整性仍具挑战。以往方法依赖多个爬虫结果对比或与外部真实数据比较。本文提出仅使用存档自身的纵向数据(即多次后续抓取的数据)来估计爬取的绝对覆盖率。核心思路是:通过连续抓取间的实际网页重叠,利用简单瓮模型拟合,再用线性回归推断模型参数。应用于德国学术网络的聚焦抓取(2013-2021年共15次半年度爬取),在稳定抓取阶段估算出覆盖率为约46%的可爬取网址空间。该方法极为简单,无需外部真实数据,可推广至任意纵向聚焦抓取场景。
原文摘要 · Abstract (English)
Web archives preserve portions of the web, but quantifying their completeness remains challenging. Prior approaches have estimated the coverage of a crawl by either comparing the outcomes of multiple crawlers, or by comparing the results of a single crawl to external ground truth datasets. We propose a method to estimate the absolute coverage of a crawl using only the archive's own longitudinal data, i.e., the data collected by multiple subsequent crawls. Our key insight is that coverage can be estimated from the empirical URL overlaps between subsequent crawls, which are in turn well described by a simple urn process. The parameters of the urn model can then be inferred from longitudinal crawl data using linear regression. Applied to our focused crawl configuration of the German Academic Web, with 15 semi-annual crawls between 2013-2021, we find a coverage of approximately 46 percent of the crawlable URL space for the stable crawl configuration regime. Our method is extremely simple, requires no external ground truth, and generalizes to any longitudinal focused crawl.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。