arXiv:2608.16390cs.CLcs.AI2026-08

Web-PDF语料统计单位混乱,导致文本损失被严重低估。

Counting Documents Is Not Counting Text: Unit Bias in Web-PDF Corpus Statistics

  • 以文档为单位统计,但实际文本分布极不均衡
  • 3.02%的文档贡献了50%的文本,1.66%的TeX文档占4.05%文本
  • 现行截断策略使55%-62%的文本永久丢失,需双单位报告

PDF语料库虽以词元(tokens)规模宣传,但其覆盖率、OCR路由、重抓取恢复率、语言混合率等指标均按文档计算,且未分解总词元数。两者单位差异显著。在包含790万网页PDF、326亿词元的CC-MAIN-2021-31-PDF-UNTRUNCATED数据集中,仅3.02%的含文本文档承载了半数词元(基尼系数0.807);超过50页的文档占总文档数的5.00%,却包含53.53%的文本。由TeX工具链生成的PDF仅占1.66%文档量,却占4.05%文本。最严重的是Common Crawl的截断上限:影响23.06%文档,却导致63.08%文本丢失。两个主流库对截断文件的恢复率分别为11.4%和1.4%,72%-97%受影响文档无恢复可能,约55%-62%的语料文本已丢失。若采用2025年3月起的5 MiB上限,仍会有30.19%词元被截断,恢复率仅从3.3%升至13.2%。建议同时报告文档与词元维度的统计结果。

原文摘要 · Abstract (English)

PDF corpora advertise their size in tokens but compute every rate they publish (coverage, OCR routing, re-fetch recovery, language mix) per document, and none decomposes its token total. The two units diverge sharply. On CC-MAIN-2021-31-PDF-UNTRUNCATED (7.9M web PDFs, 32.6B tokens), 3.02% of text-bearing documents hold half the tokens (Gini 0.807); documents over 50 pages are 5.00% of the corpus but 53.53% of its text. The PDFs produced by a TeX{} toolchain are 1.66% of documents and 4.05% of the text. The clearest casualty is Common Crawl's truncation cap: it affected 23.06% of documents and 63.08% of the text. Reconstructing the truncated files and extracting both versions, two widely used libraries recover 11.4% and 1.4% of that text; between 72% and 97% of affected documents yield nothing; roughly 55--62% of the corpus's text is lost. Under the 5 MiB cap adopted in March 2025, 30.19% of tokens would still be truncated, and recovery on those documents rises only from 3.3% to 13.2%. We recommend that corpus statistics be reported in both units: documents and tokens.

数据集分析文本丢失统计偏差

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。