arXiv:2503.02358cs.CVcs.AI2025-03ICLR被引 20

用游戏测试大模型的视觉推理能力,发现其短板。

Are Large Vision Language Models Good Game Players?

  • 设计游戏评估框架,分四类任务测视觉感知与推理。
  • 模型在长输出和密集细节识别上表现不佳。
  • 适合研究多轮决策与复杂环境理解的学者参考。

大型视觉语言模型(LVLMs)在理解和推理视觉与文本信息方面展现出显著能力。然而,现有的评估方法主要依赖于如视觉问答和图像描述等基准,常无法全面反映LVLMs的真实能力。这些基准存在视觉感知不充分、数据污染及缺乏多轮推理关注等问题。为此,我们提出 extmethod{},一个基于游戏的评估框架,旨在系统评估LVLMs在结构化环境中的认知与推理能力。该框架通过一系列游戏,评估模型在感知、问答、规则遵循和端到端游戏四个核心任务上的表现,分别考察视觉感知、推理、决策等能力。基于此框架,我们开展了大量实验,揭示了当前LVLMs在处理长结构化输出及识别详细密集元素方面的局限性。代码与数据已公开于https://github.com/xinke-wang/LVLM-Playground。

原文摘要 · Abstract (English)

Large Vision Language Models (LVLMs) have demonstrated remarkable abilities in understanding and reasoning about both visual and textual information. However, existing evaluation methods for LVLMs, primarily based on benchmarks like Visual Question Answering and image captioning, often fail to capture the full scope of LVLMs' capabilities. These benchmarks are limited by issues such as inadequate assessment of detailed visual perception, data contamination, and a lack of focus on multi-turn reasoning. To address these challenges, we propose \method{}, a game-based evaluation framework designed to provide a comprehensive assessment of LVLMs' cognitive and reasoning skills in structured environments. \method{} uses a set of games to evaluate LVLMs on four core tasks: Perceiving, Question Answering, Rule Following, and End-to-End Playing, with each target task designed to assess specific abilities, including visual perception, reasoning, decision-making, etc. Based on this framework, we conduct extensive experiments that explore the limitations of current LVLMs, such as handling long structured outputs and perceiving detailed and dense elements. Code and data are publicly available at https://github.com/xinke-wang/LVLM-Playground.

视觉语言模型游戏评估多轮推理

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。