arXiv:2510.12750cs.CVcs.AI2025-10ICCV被引 6

构建艺术与文化遗产领域高语义VQA基准,评估模型深层视觉理解能力

VQArt-Bench: A semantically rich VQA Benchmark for Art and Cultural Heritage

  • 通过多智能体协作生成复杂语义问题,提升题目多样性与深度
  • 14个顶尖模型在该基准上表现不佳,甚至简单计数任务也出错
  • 适合研究艺术理解、视觉推理及开放模型评估的学者使用

多模态大语言模型在图文联合任务中展现出强大能力,但现有视觉问答(VQA)基准难以评估深层语义理解,尤其在视觉艺术分析等复杂领域。现有题目多限于简单句式和表层属性,无法体现人类视觉探究的多样性和深度,导致模型依赖统计捷径而非真实视觉推理。为弥补这一缺陷,我们提出VQArt-Bench,一个面向文化遗产领域的大型、高语义丰富度的VQA基准。该基准采用新型多智能体流程,由专业智能体协作生成语义细腻、经验证且语言多样的问题。基准结构涵盖视觉理解的关键维度,考察模型对象征意义、叙事内容及复杂视觉关系的解析能力。对14个主流多模态大模型的评估显示,当前模型存在显著局限,包括在简单计数任务上的意外薄弱,以及专有模型与开源模型间的明显性能差距。

原文摘要 · Abstract (English)

Multimodal Large Language Models (MLLMs) have demonstrated significant capabilities in joint visual and linguistic tasks. However, existing Visual Question Answering (VQA) benchmarks often fail to evaluate deep semantic understanding, particularly in complex domains like visual art analysis. Confined to simple syntactic structures and surface-level attributes, these questions fail to capture the diversity and depth of human visual inquiry. This limitation incentivizes models to exploit statistical shortcuts rather than engage in visual reasoning. To address this gap, we introduce VQArt-Bench, a new, large-scale VQA benchmark for the cultural heritage domain. This benchmark is constructed using a novel multi-agent pipeline where specialized agents collaborate to generate nuanced, validated, and linguistically diverse questions. The resulting benchmark is structured along relevant visual understanding dimensions that probe a model's ability to interpret symbolic meaning, narratives, and complex visual relationships. Our evaluation of 14 state-of-the-art MLLMs on this benchmark reveals significant limitations in current models, including a surprising weakness in simple counting tasks and a clear performance gap between proprietary and open-source models.

视觉问答艺术理解多模态评估基准

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。