为古希腊陶器设计智能数字博物馆,让AI回答更可信、有依据。
VaseMuseum: Digital Intelligent Museum for Ancient Greek Pottery

- 构建多模态代理系统,结合3D感知与权威知识检索。
- 通过证据筛选和响应校验,减少幻觉,提升答案可靠性。
- 无需训练即可增强AI的可信度,适合文化遗产领域应用。
视觉语言模型(VLMs)使交互式数字博物馆成为可能,将3D数字化与自然语言文物探索相连接。然而,在古希腊陶器等文化遗产领域,可靠的VLM支持受限于两大挑战:其一,开放式解读需将细粒度2D/3D视觉证据与专业的策展知识对齐,但检索过程可能引入弱来源和不可验证的参考;其二,当可用证据不完整、嘈杂或模糊时,VLM常生成自信但无依据的答案。为此,我们提出VaseMuseum,一个轻量级、模块化的多模态智能代理框架,用于古希腊陶器的智能数字博物馆。VaseMuseum结合交互式虚拟博物馆与VaseAgent,支持2D图像和3D文物的多模态感知、3D-aware推理、外部知识检索及推理时可靠性控制。具体而言,VaseAgent从权威网络和博物馆知识源中检索证据,并通过源级控制选择多样且可验证的证据再生成。同时,响应级控制检查生成内容是否匹配证据池,当支持不足或冲突时,鼓励中立、基于证据的回答。此外,一种免训练的GRPO风格选择机制,通过偏好选择具备有效引用和校准置信度的响应,无需更新VLM主干。在真实数字博物馆模拟实验中,VaseMuseum显著提升了引文有效性,减少了知识密集型查询中的幻觉,并在模糊情境下产生更中立的回答,优于带搜索功能的VLM基线。
原文摘要 · Abstract (English)
Vision-language models (VLMs) have made interactive digital museums increasingly feasible by connecting 3D digitization with natural-language artifact exploration. However, in cultural heritage domains such as ancient Greek pottery, reliable VLM assistance is limited by two challenges. First, open-ended interpretation requires grounding fine-grained 2D/3D visual evidence in specialized curatorial knowledge, yet the retrieval process may introduce weak sources and unverifiable references. Second, when the available evidence is incomplete, noisy, or ambiguous, VLMs often produce confident but unsupported answers instead of calibrated uncertainty. To address these challenges, we propose VaseMuseum, a lightweight and modular multimodal agent framework for intelligent digital museums of ancient Greek pottery. VaseMuseum combines an interactive virtual museum with VaseAgent, which supports both 2D images and 3D artifacts through multimodal perception, 3D-aware reasoning, external knowledge retrieval, and inference-time reliability control. Specifically, VaseAgent retrieves evidence from authoritative web and museum knowledge sources, and source-level control selects diverse and verifiable evidence before generation. Meanwhile, response-level control checks generated claims against the evidence pool and encourages neutral, evidence-bounded answers when support is insufficient or conflicting. Moreover, a training-free GRPO-style selection mechanism favors responses with valid references and calibrated confidence without updating the VLM backbone. Experiments in a realistic digital museum simulation show that VaseMuseum improves citation validity, reduces hallucinations on knowledge-intensive queries, and produces more neutral answers under ambiguity compared with search-enabled VLM baselines.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。