arXiv:2608.19026cs.CLcs.DL2026-08

为古籍数字化文本提供可定制的多语言清洗与标注工具,保留原始信息并支持灵活使用。

Institutional Books - Enriched Text: A customizable multilingual open-source pipeline for denoising, deduplicating, and annotating OCR text at scale

  • 通过分段语言识别和去重聚类,保留元数据的同时清理文本
  • 产出1.39亿个带标注的段落,涵盖约250种语言,总词元量2170亿
  • 适合需要精准处理历史文献的研究者与开发者,避免重复劳动

2025年发布的哈佛图书馆典藏(IB-HL)包含983,004册图书(2420亿o200k_base词元),源自谷歌图书计划。随着研究者使用该数据集,标准预处理流程与信息负责任管理之间产生矛盾:现有流水线常过度过滤、去重、限制语言,甚至丢弃重要元数据。为解决此问题,我们提出Enriched Text方法——不生成单一文本流,而是通过注解层保留元数据,在保留原文基础上完成分段语言识别、重复段落聚类及每段比特/字节评分。所有结果以类似HTML的注解形式附加于文本之上,用户可按需解析。该流程覆盖约250种语言。本文发布包含2170亿词元、1.39亿个带注释子主题段落的IB-HL-ET版本及其完整处理流水线,提升机器可读性与人工可研性。

原文摘要 · Abstract (English)

Released in 2025, Institutional Books: Harvard Library (IB-HL) is a collection of 983,004 volumes (242B o200k_base tokens), originally digitized through Harvard Library's participation in the Google Books Library project. As researchers and developers have begun to use IB-HL, a tension has emerged between standard large-scale preprocessing practices and the goals of careful information stewardship. Many existing pipelines optimize for web text: as a result, they tend to aggressively filter, deduplicate, restrict by language, and sometimes discard meaningful metadata. Meanwhile, researchers seeking to use IB-HL duplicate effort while performing similar processing and analysis. We describe an approach that we call Enriched Text. Instead of producing a single 'complete' stream of tokens, we normalize the text while preserving metadata through annotations. We separate endmatter, detect per-paragraph language, identify clusters of duplicate paragraphs, and compute per-paragraph bits-per-byte scores. We provide this information through HTML-like annotations layered on top of the text. By parsing these annotations, users can tailor the output to their own needs instead of accepting a global editorial decision on content. The pipeline applies to all $\approx$250 languages in the collection. This report describes this project's goals, implementation, and design rationale. The release includes IB-HL-ET (an enriched-text version of IB-HL containing 217B o200k_base tokens across 983,003 volumes, organized into 1.39B annotated subtopic paragraphs) and the pipeline that produced it. These serve to make the collection easier for machines to parse and for humans to study.

文本清洗多语言元数据开放数据

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。