构建古希腊陶器视觉问答数据集,提升AI对文化遗产的深度理解能力。
VaseVQA: Multimodal Agent and Benchmark for Ancient Greek Pottery
- 基于7类专家定义类别构建3万+图像与6万+问答对数据集
- 引入可验证奖励的强化学习策略,显著提升复杂推理准确率
- 适合研究文化遗产理解、多模态推理与领域适应的学者使用
理解古希腊陶器等文化遗产需要专家级推理能力,而当前多模态大模型因缺乏特定领域数据难以胜任。本文提出VaseVQA,包含31,773张图像和67,614个问答对,覆盖七类专家定义类别,支持对文化遗产理解能力的系统评估。基于该数据集,我们探索了特定领域的训练策略。虽然监督微调能提升领域知识适配性,但在深层推理任务上表现有限。为此,我们提出VaseVL,通过可验证奖励的强化学习增强监督微调。实验表明,VaseVL在推理密集型问题上持续优于监督基线,凸显针对性强化学习在文化遗产视觉问答中的价值。代码与数据集将开源于https://github.com/AIGeeksGroup/VaseVQA。
原文摘要 · Abstract (English)
Understanding cultural heritage artifacts such as ancient Greek pottery requires expert-level reasoning that remains challenging for current MLLMs due to limited domain-specific data. We introduce VaseVQA, a benchmark of 31,773 images and 67,614 question-answer pairs across seven expert-defined categories, enabling systematic evaluation of expert-level cultural heritage understanding. Using this dataset, we explore effective training strategies for domain-specific reasoning. While supervised fine-tuning improves adaptation to domain knowledge, it struggles with deeper reasoning tasks. We propose VaseVL, which augments SFT with reinforcement learning using verifiable rewards. Experiments show that VaseVL consistently outperforms supervised baselines, especially on reasoning-intensive questions, highlighting the value of targeted reinforcement learning for cultural heritage visual question answering. Our code and dataset will be released at https://github.com/AIGeeksGroup/VaseVQA.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。