arXiv:2411.04663cs.CV2024-11被引 6

用多模态大模型实现可解释的文化遗产图像搜索与发现

Explainable Search and Discovery of Visual Cultural Heritage Collections with Multimodal Large Language Models

  • 基于多模态大模型构建开放可解释的图像搜索接口
  • 生成推荐结果并提供具体文本说明,无需预先设定特征
  • 适合文化遗产数字化、数字人文研究者使用

许多文化机构已将大规模数字化视觉资料在线公开,通常允许再利用。但缺乏细粒度元数据使得创建探索和搜索界面极具挑战。本文提出一种利用先进多模态大语言模型的方法,实现对视觉藏品的开放式、可解释的搜索与发现。我们展示了该方法如何构建新颖的聚类与推荐系统,避免直接依赖视觉嵌入带来的常见问题。特别值得关注的是,系统能为每项推荐提供具体的文本解释,无需预先选定关注特征。这些特性共同打造了一个更开放灵活的数字界面,同时更契合隐私与伦理要求。通过纪录片照片收藏的案例研究,我们提供了多项指标,验证了该方法的有效性与潜力。

原文摘要 · Abstract (English)

Many cultural institutions have made large digitized visual collections available online, often under permissible re-use licences. Creating interfaces for exploring and searching these collections is difficult, particularly in the absence of granular metadata. In this paper, we introduce a method for using state-of-the-art multimodal large language models (LLMs) to enable an open-ended, explainable search and discovery interface for visual collections. We show how our approach can create novel clustering and recommendation systems that avoid common pitfalls of methods based directly on visual embeddings. Of particular interest is the ability to offer concrete textual explanations of each recommendation without the need to preselect the features of interest. Together, these features can create a digital interface that is more open-ended and flexible while also being better suited to addressing privacy and ethical concerns. Through a case study using a collection of documentary photographs, we provide several metrics showing the efficacy and possibilities of our approach.

文化遗产多模态可解释性大模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。