arXiv:2510.04479cs.CV2025-10被引 5

首个针对古希腊陶器的3D视觉问答数据集,提升AI对文物的理解能力。

VaseVQA-3D: Benchmarking 3D VLMs on Ancient Greek Pottery

  • 构建首个古希腊陶器3D VQA数据集,含664个模型及问答对。
  • 新模型在R@1指标上提升12.8%,词汇相似度提高6.6%。
  • 适合数字文化遗产保护与跨模态文物分析研究者使用。

视觉语言模型(VLMs)在多模态理解任务中取得显著进展,尤其在图像描述和视觉推理等通用任务中表现突出。然而,在古希腊陶器等专业文化遗产领域,现有模型面临严重数据稀缺和领域知识不足的问题。由于缺乏针对性训练数据,当前VLMs难以有效处理此类文化重要任务。为此,我们提出VaseVQA-3D数据集,作为首个用于古希腊陶器分析的3D视觉问答数据集,收集了664个古希腊陶器3D模型及其对应的问答数据,并建立了完整的数据构建流程。我们进一步开发了VaseVLM模型,通过领域自适应训练提升模型在陶器分析中的性能。实验结果验证了方法的有效性:相比先前最优模型,在VaseVQA-3D数据集上R@1指标提升12.8%,词汇相似度提升6.6%,显著增强了对3D陶器文物的识别与理解能力,为数字遗产保护研究提供了新的技术路径。

原文摘要 · Abstract (English)

Vision-Language Models (VLMs) have achieved significant progress in multimodal understanding tasks, demonstrating strong capabilities particularly in general tasks such as image captioning and visual reasoning. However, when dealing with specialized cultural heritage domains like 3D vase artifacts, existing models face severe data scarcity issues and insufficient domain knowledge limitations. Due to the lack of targeted training data, current VLMs struggle to effectively handle such culturally significant specialized tasks. To address these challenges, we propose the VaseVQA-3D dataset, which serves as the first 3D visual question answering dataset for ancient Greek pottery analysis, collecting 664 ancient Greek vase 3D models with corresponding question-answer data and establishing a complete data construction pipeline. We further develop the VaseVLM model, enhancing model performance in vase artifact analysis through domain-adaptive training. Experimental results validate the effectiveness of our approach, where we improve by 12.8% on R@1 metrics and by 6.6% on lexical similarity compared with previous state-of-the-art on the VaseVQA-3D dataset, significantly improving the recognition and understanding of 3D vase artifacts, providing new technical pathways for digital heritage preservation research. Code: https://github.com/AIGeeksGroup/VaseVQA-3D. Website: https://aigeeksgroup.github.io/VaseVQA-3D.

3D视觉问答文物分析多模态数字遗产

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。