arXiv:2511.20272cs.CV2025-11中稿 · ECCV被引 5

评测多模态大模型对视觉常识的理解能力,发现现有模型仍远不如人类。

VKnowU: Evaluating Visual Knowledge Understanding in Multimodal LLMs

  • 构建涵盖1249个视频的评测集VKnowU,覆盖8类视觉常识
  • 顶尖模型在物理常识任务上表现仅达人类水平的67.3%(3.7%提升空间)
  • 提出VideoKnow+模型,通过强化学习提升模型对视觉常识的理解

尽管多模态大语言模型已擅长物体识别,但缺乏对世界深层物理与社会规则的直观理解。我们称之为视觉知识的高层语义,是连接感知与推理的桥梁,但当前研究尚未充分探索。为此,我们提出VKnowU,一个包含1,680个问题、1,249个视频的综合性评测基准,涵盖8类核心视觉知识,包括世界中心型(如直觉物理)和人类中心型(如主观意图)。对28个前沿多模态大模型的评估显示,领先模型仍显著落后于人类表现,尤其在世界中心型任务上差距明显。为弥补这一差距,我们引入新数据集VKnowQA及基线模型VideoKnow+,该模型采用结构化‘看-想-答’范式,并结合视觉知识奖励的强化学习,使在VKnowU上提升3.7%,并在MVBench(+5.4%)、Video-MME(+7.0%)和MMVU(+5.7%)上均实现持续增益。本工作强调视觉知识是构建更泛化多模态大模型的关键基石,使其不仅能‘看’,更能‘理解’真实世界。

原文摘要 · Abstract (English)

While Multimodal Large Language Models (MLLMs) have become adept at recognizing objects, they often lack the intuitive, human-like understanding of the world's underlying physical and social principles. This high-level vision-grounded semantics, which we term visual knowledge, forms a bridge between perception and reasoning, yet remains an underexplored area in current MLLMs. To systematically evaluate this capability, we present VKnowU, a comprehensive benchmark featuring 1,680 questions in 1,249 videos, covering 8 core types of visual knowledge spanning both world-centric (e.g., intuitive physics) and human-centric (e.g., subjective intentions). Evaluation of 28 SOTA MLLMs reveals that leading models still fall short of human performance, with particularly notable gaps in the world-centric. To bridge this gap, we introduce a new dataset, VKnowQA, and VideoKnow+, a baseline model that explicitly incorporates visual knowledge into MLLMs. VideoKnow+ follows a structured See-Think-Answer paradigm and adopts reinforcement learning with visual knowledge reward, achieving a +3.7% improvement on VKnowU and consistent gains on MVBench (+5.4%), Video-MME (+7.0%), and MMVU (+5.7%). Our work highlights visual knowledge as a missing cornerstone for developing more generalizable MLLMs that can not only see but also truly understand our worlds.

多模态视觉常识评测基准强化学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。