解决图书馆从Google Books下载数据时的速率限制和元数据碎片问题。
GRIN Transfer: A production-ready tool for libraries to retrieve digital copies from Google Books
- 开发了Python工具GRIN Transfer,支持高效批量下载
- 可处理数百万条扫描件数据,克服平台速率限制
- 适合需要结构化获取Google Books数据的图书馆
自2004年上线以来,Google Books项目已与全球图书馆合作扫描超千万册图书。通过其回传接口(GRIN),图书馆可访问扫描件、元数据及经新技术重处理后的改进数据。在构建机构图书数据集过程中,我们发现GRIN存在速率限制和元数据碎片化问题。为此,本文发布GRIN Transfer——一个开源、生产就绪的Python数据提取管道,帮助合作图书馆高效获取其谷歌图书收藏。同时更新了机构图书1.0数据处理流程,确保与GRIN Transfer输出格式兼容。二者结合可形成从下载、结构化到增强的完整处理链路。本报告详述了该工具在不同环境下的可靠性设计与配置指南。
原文摘要 · Abstract (English)
Publicly launched in 2004, the Google Books project has scanned tens of millions of items in partnership with libraries around the world. As part of this project, Google created the Google Return Interface (GRIN). Through this platform, libraries can access their scanned collections, the associated metadata, and the ongoing OCR and metadata improvements that become available as Google reprocesses these collections using new technologies. When downloading the Harvard Library Google Books collection from GRIN to develop the Institutional Books dataset, we encountered several challenges related to rate-limiting and atomized metadata within the GRIN platform. To overcome these challenges and help other libraries make more robust use of their Google Books collections, this technical report introduces the initial release of GRIN Transfer. This open-source and production-ready Python pipeline allows partner libraries to efficiently retrieve their Google Books collections from GRIN. This report also introduces an updated version of our Institutional Books 1.0 pipeline, initially used to analyze, augment, and assemble the Institutional Books 1.0 dataset. We have revised this pipeline for compatibility with the output format of GRIN Transfer. A library could pair these two tools to create an end-to-end processing pipeline for their Google Books collection to retrieve, structure, and enhance data available from GRIN. This report gives an overview of how GRIN Transfer was designed to optimize for reliability and usability in different environments, as well as guidance on configuration for various use cases.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。