用大模型和检索增强生成,自动给早期宗教木刻分类
Automating Iconclass: LLMs and RAG for Large-Scale Classification of Religious Woodcuts
- 结合大模型与向量库,利用全页图文上下文生成描述
- 在五级和四级分类上分别达到87%和92%精确率
- 适合艺术史与数字人文研究者处理大规模图像档案
本文提出一种新方法,利用大型语言模型(LLMs)与向量数据库结合检索增强生成(RAG),对早期现代宗教图像进行分类。该方法基于神圣罗马帝国书籍插图的全页上下文,使大模型能生成融合视觉与文本信息的详细描述,并通过混合向量搜索匹配相应的Iconclass代码。实验显示,该方法在五级和四级分类上分别达到87%和92%的精确率,显著优于传统图像与关键词搜索。通过使用全页描述与RAG,系统提升了分类准确性,为早期现代视觉档案的大规模分析提供了有力工具。这一跨学科方法展示了大模型与RAG在艺术史与数字人文领域的巨大潜力。
原文摘要 · Abstract (English)
This paper presents a novel methodology for classifying early modern religious images by using Large Language Models (LLMs) and vector databases in combination with Retrieval-Augmented Generation (RAG). The approach leverages the full-page context of book illustrations from the Holy Roman Empire, allowing the LLM to generate detailed descriptions that incorporate both visual and textual elements. These descriptions are then matched to relevant Iconclass codes through a hybrid vector search. This method achieves 87% and 92% precision at five and four levels of classification, significantly outperforming traditional image and keyword-based searches. By employing full-page descriptions and RAG, the system enhances classification accuracy, offering a powerful tool for large-scale analysis of early modern visual archives. This interdisciplinary approach demonstrates the growing potential of LLMs and RAG in advancing research within art history and digital humanities.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。