用游戏评测多模态大模型的视觉交互能力,发现其在复杂场景下表现明显不足。
V-MAGE: A Game Evaluation Framework for Assessing Vision-Centric Capabilities in Multimodal Large Language Models
- 设计五款视频游戏,构建动态视觉环境评估模型交互推理能力
- 领先模型在简单任务接近人类,复杂任务性能骤降超40%
- 基于动态ELO评分,可公平对比不同难度与任务类型下的模型表现
多模态大语言模型(MLLMs)在图文处理上已表现出色,但现有静态图像-文本基准无法评估其动态感知与交互推理能力。本文提出视觉中心型多能力游戏评测框架V-MAGE,通过五款视频游戏、30余个精心设计的评测场景,构建自由形式、视觉复杂的连续空间环境,要求模型仅凭视觉输入理解动态游戏状态并作出决策,贴近真实玩家体验。为实现稳健可解释的跨模型比较,V-MAGE采用动态ELO评分系统,兼顾任务难度差异与多样性。对主流MLLMs的测试表明,尽管顶尖模型在简单任务中接近人类水平,但在需高级推理与任务编排的复杂场景中性能显著下降。该性能差距揭示了当前模型在连续时间环境中进行视觉引导的逐帧交互控制方面存在根本性局限。通过深入分析,验证了V-MAGE在识别模型缺陷与指导改进方向方面的有效性。代码已开源:https://github.com/CSU-JPG/V-MAGE。
原文摘要 · Abstract (English)
Recent advancements in Multimodal Large Language Models (MLLMs) have demonstrated impressive capabilities in visual-text processing. However, existing static image-text benchmarks are insufficient for evaluating their dynamic perception and interactive reasoning abilities. We introduce Vision-centric Multiple Abilities Game Evaluation (V-MAGE), a novel game-based evaluation framework designed to systematically assess MLLMs' visual reasoning in interactive, continuous-space environments. V-MAGE features five distinct video games comprising over 30 carefully constructed evaluation scenarios. These scenarios are set in free-form, visually complex environments that require models to interpret dynamic game states and make decisions based solely on visual input, thereby closely reflecting the conditions encountered by human players. To ensure robust and interpretable comparisons across models, V-MAGE employs a dynamic ELO-based ranking system that accounts for varying difficulty levels and task diversity. Benchmarking state-of-the-art MLLMs against human baselines reveals that while leading models approach human-level performance in simple tasks, their performance drops significantly in complex scenarios requiring advanced reasoning and task orchestration. This persistent performance gap highlights fundamental limitations in current MLLMs' ability to perform vision-grounded, interactive frame-by-frame control in simulated continuous-time environments. Through extensive analyses, we demonstrate the utility of V-MAGE in uncovering these limitations and providing actionable insights for improving the visual and reasoning capabilities of MLLMs in dynamic, interactive settings. Code is publicly available at https://github.com/CSU-JPG/V-MAGE.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。