测试大模型玩视觉游戏,发现连基础空间推理都极弱
ING-VP: MLLMs cannot Play Easy Vision-based Games Yet
- 设计6种游戏共300关,每关6种配置,累计超6万轮交互
- 顶尖模型Claude-3.5 Sonnet平均准确率仅3.37%,远低于预期
- 专攻空间想象与多步推理,适合研究视觉规划的学者
随着多模态大语言模型在各类任务中表现日益出色,更复杂全面的评测基准应运而生,以检验其感知、推理与规划等核心能力。然而,现有基准在评估基于图像空间关系的多步规划能力方面仍显不足。为此,我们提出ING-VP——首个基于交互式游戏的视觉规划评测基准,专门用于评估多模态大模型的空间想象与多步推理能力。ING-VP包含6种不同游戏,共300个关卡,每关有6种独特配置,单个模型需完成超过60,000轮交互。该框架支持多种对比设置,包括图像-文本与纯文本输入、单步与多步推理、带历史与无历史条件,可深入揭示模型表现。我们评测了多个先进MMLMs,最高性能模型Claude-3.5 Sonnet平均准确率仅为3.37%,远低于预期。本工作旨在提供一个专项评测框架,推动多模态大模型在复杂空间推理与规划能力上的进步。代码已公开于https://github.com/Thisisus7/ING-VP.git。
原文摘要 · Abstract (English)
As multimodal large language models (MLLMs) continue to demonstrate increasingly competitive performance across a broad spectrum of tasks, more intricate and comprehensive benchmarks have been developed to assess these cutting-edge models. These benchmarks introduce new challenges to core capabilities such as perception, reasoning, and planning. However, existing multimodal benchmarks fall short in providing a focused evaluation of multi-step planning based on spatial relationships in images. To bridge this gap, we present ING-VP, the first INteractive Game-based Vision Planning benchmark, specifically designed to evaluate the spatial imagination and multi-step reasoning abilities of MLLMs. ING-VP features 6 distinct games, encompassing 300 levels, each with 6 unique configurations. A single model engages in over 60,000 rounds of interaction. The benchmark framework allows for multiple comparison settings, including image-text vs. text-only inputs, single-step vs. multi-step reasoning, and with-history vs. without-history conditions, offering valuable insights into the model's capabilities. We evaluated numerous state-of-the-art MLLMs, with the highest-performing model, Claude-3.5 Sonnet, achieving an average accuracy of only 3.37%, far below the anticipated standard. This work aims to provide a specialized evaluation framework to drive advancements in MLLMs' capacity for complex spatial reasoning and planning. The code is publicly available at https://github.com/Thisisus7/ING-VP.git.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。