用多模态检索增强生成系统分析文物来源,提升考古研究效率
Provenance Analysis of Archaeological Artifacts via Multimodal RAG Systems
- 构建图文双模知识库,支持视觉、边缘和语义多路检索
- 在英博馆东欧亚青铜时代文物上验证,输出具时间地理文化属性的推断
- 输出可解释,减轻学者对比海量资料的认知负担
本文提出一种基于检索增强生成(RAG)的文物来源分析系统,通过整合多模态检索与大视觉语言模型(VLMs),辅助专家推理。系统从参考文本与图像构建双模知识库,实现原始视觉、边缘增强与语义检索,识别风格相似文物。检索候选对象由VLM合成,生成包含年代、地理与文化归属的结构化推论及解释性依据。我们在英国博物馆的东欧亚青铜时代文物数据集上评估该系统,专家评价表明其输出具有意义且可解释,为学者提供明确分析起点,显著降低浏览庞大比对语料的认知负荷。
原文摘要 · Abstract (English)
In this work, we present a retrieval-augmented generation (RAG)-based system for provenance analysis of archaeological artifacts, designed to support expert reasoning by integrating multimodal retrieval and large vision-language models (VLMs). The system constructs a dual-modal knowledge base from reference texts and images, enabling raw visual, edge-enhanced, and semantic retrieval to identify stylistically similar objects. Retrieved candidates are synthesized by the VLM to generate structured inferences, including chronological, geographical, and cultural attributions, alongside interpretive justifications. We evaluate the system on a set of Eastern Eurasian Bronze Age artifacts from the British Museum. Expert evaluation demonstrates that the system produces meaningful and interpretable outputs, offering scholars concrete starting points for analysis and significantly alleviating the cognitive burden of navigating vast comparative corpora.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。