arXiv:2506.06242cs.CVcs.AI2025-06ICML被引 4

构建视觉图谱数据集,测试AI是否能识别不同形态下的相同概念。

Visual Graph Arena: Evaluating Visual Conceptualization of Vision and Multimodal Large Language Models

  • 用多种图形布局设计任务,检验模型对视觉形式无关概念的推理能力。
  • 人类准确率接近100%,模型在同构检测上完全失败,路径/环路任务表现有限。
  • 揭示模型存在伪智能模式匹配,不具真正理解,适合研究视觉抽象的学者。

多模态大模型在视觉问答中取得突破,但一个关键差距仍存:概念化——即在视觉形式变化下识别和推理同一概念的能力,这是人类推理的基本能力。为解决此问题,我们提出视觉图谱竞技场(Visual Graph Arena, VGA),包含六个基于图的任务,用于评估与提升AI在视觉抽象方面的能力。VGA采用多样化图布局(如Kamada-Kawai与平面布局)测试推理对视觉形式的独立性。实验显示,人类在所有任务中接近完美准确率,而模型在同构检测任务上完全失败,在路径与环路任务中仅表现出有限成功。我们还发现行为异常,表明模型依赖伪智能模式匹配而非真实理解。这些结果凸显当前AI在视觉理解上的根本局限。通过分离表示不变推理挑战,VGA为推动类人概念化提供框架。数据集已开放:vga.csail.mit.edu

原文摘要 · Abstract (English)

Recent advancements in multimodal large language models have driven breakthroughs in visual question answering. Yet, a critical gap persists, `conceptualization'-the ability to recognize and reason about the same concept despite variations in visual form, a basic ability of human reasoning. To address this challenge, we introduce the Visual Graph Arena (VGA), a dataset featuring six graph-based tasks designed to evaluate and improve AI systems' capacity for visual abstraction. VGA uses diverse graph layouts (e.g., Kamada-Kawai vs. planar) to test reasoning independent of visual form. Experiments with state-of-the-art vision models and multimodal LLMs reveal a striking divide: humans achieved near-perfect accuracy across tasks, while models totally failed on isomorphism detection and showed limited success in path/cycle tasks. We further identify behavioral anomalies suggesting pseudo-intelligent pattern matching rather than genuine understanding. These findings underscore fundamental limitations in current AI models for visual understanding. By isolating the challenge of representation-invariant reasoning, the VGA provides a framework to drive progress toward human-like conceptualization in AI visual models. The Visual Graph Arena is available at: \href{https://vga.csail.mit.edu/}{vga.csail.mit.edu}

视觉理解概念化多模态模型图神经网络

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。