arXiv:2512.10165cs.DLcs.IR2025-12

自动为书目数据补全标识符并聚类同一作品的不同版本。

BookReconciler: An Open-Source Tool for Metadata Enrichment and Work-Level Clustering

  • 基于OpenRefine插件,自动匹配权威标识符如ISBN。
  • 对同一作品的译本、版本进行聚类,准确率达98%(美国获奖作品)。
  • 适合数字人文与图书馆学研究者,提升跨库书目整合效率。

我们提出BookReconciler,一个开源工具,用于增强和聚类图书数据。用户只需提供包含书名和作者的简易表格,该工具即可自动添加如ISBN等权威持久标识符,并将同一作品的不同表达形式(如不同译本或版本)聚类。该功能便于合并相关文献并实现大规模分析。工具目前作为OpenRefine扩展,对接美国国会图书馆、VIAF、OCLC、HathiTrust、Google Books和Wikidata等主流书目服务。方法强调人工判断,通过交互界面允许用户评估匹配结果并界定作品边界(例如是否包含译本)。在美籍获奖作品和当代世界文学数据集上评估显示,对美国作品的重合准确率接近完美,但全球文本表现较低,反映出非英语及全球文学书目基础设施的结构性缺陷。整体而言,BookReconciler支持跨领域、跨应用的书目数据复用,助力数字图书馆与数字人文研究。

原文摘要 · Abstract (English)

We present BookReconciler, an open-source tool for enhancing and clustering book data. BookReconciler allows users to take spreadsheets with minimal metadata, such as book title and author, and automatically 1) add authoritative, persistent identifiers like ISBNs 2) and cluster related Expressions and Manifestations of the same Work, e.g., different translations or editions. This enhancement makes it easier to combine related collections and analyze books at scale. The tool is currently designed as an extension for OpenRefine -- a popular software application -- and connects to major bibliographic services including the Library of Congress, VIAF, OCLC, HathiTrust, Google Books, and Wikidata. Our approach prioritizes human judgment. Through an interactive interface, users can manually evaluate matches and define the contours of a Work (e.g., to include translations or not). We evaluate reconciliation performance on datasets of U.S. prize-winning books and contemporary world fiction. BookReconciler achieves near-perfect accuracy for U.S. works but lower performance for global texts, reflecting structural weaknesses in bibliographic infrastructures for non-English and global literature. Overall, BookReconciler supports the reuse of bibliographic data across domains and applications, contributing to ongoing work in digital libraries and digital humanities.

书目数据聚类数字人文

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。