通过发现曲线分析网页爬虫数据,揭示了核心与外壳的动态差异。
Measuring What the Crawler Sees: Discovery Curves, Core Persistence, and Shell Dynamics in Longitudinal Web Crawls

- 引入发现曲线,用滑动窗口追踪连续爬取的网页覆盖范围。
- 发现核心留存率κ与外壳覆盖率不一致,表明网络结构分层存在。
- 为长期爬虫分析提供可量化的多组件模型框架,适合数据科学家参考。
纵向网页爬取是随时间变化的网址集合的部分采样序列。传统方法通过两轮爬取的包含关系来评估,基于简单的抽样-替换瓮模型(urn model),可估算出每轮存活率α和覆盖率c,但假设网址分布均匀且每次仅分析一对。本文提出一套形式化语言描述爬取过程,并引入发现曲线U(s, T),即从第s轮开始、跨度T轮的累积网址覆盖范围,在相同瓮模型下可解析为(α, c)的闭式函数。容器性与发现曲线是同一过程的两个投影;若瓮模型均匀,两者独立拟合应得一致结果,任何偏差即为可观测信号。在Common Crawl(2020–2025,域名粒度)和德国学术网(GAW,URL粒度)数据上,二者均出现分歧,采用双组分瓮模型(持久核心占比κ,外壳参数(α_∂, c_∂))可调和差异。剩余的c_∂残差表明外壳本身非均匀;κ作为秩解析推广的标量入口,留待后续研究。
原文摘要 · Abstract (English)
A longitudinal web crawl is a sequence of partial samples of an evolving URL population. Pairwise containment between two crawls is the standard probe; under a simple \emph{urn} model of the crawl -- each round samples a fraction of the URLs and replaces a fraction -- it recovers two interpretable rates, per-round survival $α$ and coverage $c$, but treats the population as uniform and consumes one pair at a time. In this work, we define a formal language for talking about a crawl. We extend this analysis with the \emph{discovery curve} $U(s, T)$, the cumulative URL footprint over a sliding window of $T$ crawls starting at $s$, which under the same urn model is also a closed-form function of $(α, c)$. Containment and the discovery curve are then two projections of one process: independent fits agree on $(α, c)$ when the urn is homogeneous, so any disagreement is itself a measurement. Applied to Common Crawl (2020--2025, domain granularity) and to the German Academic Web (GAW, URL granularity), the two projections disagree on both archives, and a two-component urn with a persistent core fraction $κ$ alongside shell parameters $(α_\partial, c_\partial)$ reconciles the disagreement. A residual on $c_\partial$ remains, signaling that the shell itself is not homogeneous; $κ$ is recorded as the scalar entry point to a rank-resolved generalization, which is left to follow-up work. \keywords{web archive \and crawl coverage \and discovery curve \and urn model \and two-component model \and URL lifetime}
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。