arXiv:2608.18957cs.CVcs.DL2026-08

开源工具链批量提取并标注古籍中的图像元素,助力数字人文研究。

Institutional Books - Visual Elements: An open-source pipeline for extracting, classifying, deduplicating, and captioning visual elements from digital book collections

  • 构建端到端流程自动检测分类历史书籍中的插图、照片等视觉元素。
  • 从98.3万册扫描图书中提取2260万条视觉元素,形成首个大规模公开数据集。
  • 适合数字人文、AI训练与图书馆数字化项目使用。

历史书籍藏有丰富的视觉元素,如插图、照片、雕刻和装饰艺术,但在大规模数字化项目中常被忽视。尽管光学字符识别(OCR)已标准化文本内容提取,但这些视觉成分所蕴含的细微语境仍未能通过自动化文本流程充分挖掘。本文介绍Institutional Books - Visual Elements:一个开源的端到端处理流程,用于从历史书籍中检测、分类、去重并为视觉元素生成描述。伴随该流程,我们发布了首个数据集,包含从983,004册扫描图书中提取的2260万条视觉元素,覆盖Institutional Books: Harvard Library数据集。本工作推动了社区对数字化馆藏的计算化利用,支持人工智能模型训练与数字人文研究等新应用场景。

原文摘要 · Abstract (English)

Historical book collections contain rich visual elements - such as illustrations, photographs, engravings, and decorative art - that are frequently under-explored in large-scale digitization projects. While Optical Character Recognition (OCR) has standardized the extraction of textual content, these visual components offer a layer of nuance and context that remains largely untapped by automated text extraction workflows. This technical report introduces Institutional Books - Visual Elements, an open-source end-to-end pipeline for detecting, classifying, deduplicating, and captioning visual elements from historical book collections. Alongside this pipeline, we release an initial dataset of 22.6 million visual elements extracted from the 983,004 scanned volumes that comprise the Institutional Books: Harvard Library dataset. This work contributes to ongoing, community-wide efforts to enable new use cases for digitized library collections through computational access, from artificial intelligence model training to digital humanities research.

图像分析数字人文开源工具

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。