arXiv:2506.08300cs.CLcs.DL2025-06被引 7

2420亿词的哈佛图书馆古籍数据集,提升LLM训练数据质量与可追溯性。

Institutional Books 1.0: A 242B token dataset from Harvard Library's collections, refined for accuracy and usability

  • 从哈佛图书馆107万册古籍中筛选出2420亿词的公共领域文本
  • 覆盖250多种语言,经校对和元数据标注,可用性高
  • 适合研究历史语言、AI训练数据伦理及可追溯性的人群

大型语言模型依赖数据学习世界知识以生成有意义的关联与预测。训练数据的规模、质量、多样性及来源清晰度直接影响模型表现。随着大模型快速演进,高质量公开训练数据稀缺问题凸显,亟需建立可持续、有明确来源链的数据管理体系。本文介绍机构书籍1.0(Institutional Books 1.0),基于哈佛图书馆参与谷歌图书项目自2006年起扫描的1075899册文献,从中提取并处理出一个全面标注的历史文本数据集。原始数据涵盖超过250种语言,总计约2500亿词。此次发布包含983,004册公共领域书籍的OCR文本(原始与后处理版)及元数据(书目、来源与生成信息),总规模达2420亿词。本报告详述项目目标、方法及分析结果,旨在使这一历史文献集合更易被人类与机器过滤、阅读与使用。

原文摘要 · Abstract (English)

Large language models (LLMs) use data to learn about the world in order to produce meaningful correlations and predictions. As such, the nature, scale, quality, and diversity of the datasets used to train these models, or to support their work at inference time, have a direct impact on their quality. The rapid development and adoption of LLMs of varying quality has brought into focus the scarcity of publicly available, high-quality training data and revealed an urgent need to ground the stewardship of these datasets in sustainable practices with clear provenance chains. To that end, this technical report introduces Institutional Books 1.0, a large collection of public domain books originally digitized through Harvard Library's participation in the Google Books project, beginning in 2006. Working with Harvard Library, we extracted, analyzed, and processed these volumes into an extensively-documented dataset of historic texts. This analysis covers the entirety of Harvard Library's collection scanned as part of that project, originally spanning 1,075,899 volumes written in over 250 different languages for a total of approximately 250 billion tokens. As part of this initial release, the OCR-extracted text (original and post-processed) as well as the metadata (bibliographic, source, and generated) of the 983,004 volumes, or 242B tokens, identified as being in the public domain have been made available. This report describes this project's goals and methods as well as the results of the analyses we performed, all in service of making this historical collection more accessible and easier for humans and machines alike to filter, read and use.

数据集历史文本LLM训练

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。