arXiv:2410.14969cs.CVcs.IR2024-10中稿 · the 2024 Computati…被引 2

用视觉模型提升挪威国家图书馆古籍图像检索与分类效果

Visual Navigation of Digital Libraries: Retrieval and Classification of Images in the National Library of Norway's Digitised Book Collection

  • 对比ViT、CLIP、SigLIP三种视觉模型进行图像检索与分类
  • SigLIP在检索和分类任务中表现最优,略胜于CLIP和ViT
  • 可辅助清理数字化流程中的图像数据集,适合数字人文研究者

数字化图书馆的文本分析工具已广泛应用,而计算机视觉的进步为视觉内容分析带来新可能。由于许多古籍同时包含文字与图像,利用深度学习嵌入技术提升图像可访问性至关重要。本文针对挪威国家图书馆1900年前的古籍图像,构建了一个图像检索原型系统,比较了视觉Transformer(ViT)、对比语言-图像预训练(CLIP)以及基于Sigmoid损失的语言-图像预训练(SigLIP)三种嵌入方法在图像检索与分类中的表现。结果表明,该应用在精确图像检索上表现良好,其中SigLIP嵌入在检索与分类任务中均略优于CLIP与ViT。此外,基于SigLIP的图像分类能有效辅助清理数字化流程中的图像数据集。

原文摘要 · Abstract (English)

Digital tools for text analysis have long been essential for the searchability and accessibility of digitised library collections. Recent computer vision advances have introduced similar capabilities for visual materials, with deep learning-based embeddings showing promise for analysing visual heritage. Given that many books feature visuals in addition to text, taking advantage of these breakthroughs is critical to making library collections open and accessible. In this work, we present a proof-of-concept image search application for exploring images in the National Library of Norway's pre-1900 books, comparing Vision Transformer (ViT), Contrastive Language-Image Pre-training (CLIP), and Sigmoid loss for Language-Image Pre-training (SigLIP) embeddings for image retrieval and classification. Our results show that the application performs well for exact image retrieval, with SigLIP embeddings slightly outperforming CLIP and ViT in both retrieval and classification tasks. Additionally, SigLIP-based image classification can aid in cleaning image datasets from a digitisation pipeline.

图像检索视觉模型数字人文

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。