arXiv:2605.11223cs.AI2026-05

测试视觉语言模型在点按谜题中的类人逻辑解题能力

Do Vision-Language-Models show human-like logical problem-solving capability in point and click puzzle games?

论文配图:Do Vision-Language-Models show human-like logical problem-solving capability in point and click puzzle games?
图 1 · 摘自论文原文
  • 构建新基准VLATIM,评估模型在物理谜题中的逻辑推理与精准操作能力
  • 大模型虽有优秀规划能力,但难以准确定位与执行动作
  • 适合研究多模态交互、具身智能与人机协作的学者参考

视觉-语言-动作模型(VLMs)在交互环境中的应用日益广泛,但现有评测基准常忽视点按谜题中所需的复杂物理推理。本文提出视觉-语言对抗不可思议机器(VLATIM)基准,用于评估经典物理谜题游戏《不可思议机器2》(TIM)中的类人逻辑解题能力。该基准分为五个递进部分,涵盖从基础视觉定位到多步操作与完整解谜的多种能力。结果表明,尽管大型专有模型具备较强的规划能力,但在精确视觉定位方面表现不佳,整体尚未展现出类人问题解决能力。

原文摘要 · Abstract (English)

Vision-Language(-Action) Models (VLMs) are increasingly applied to interactive environments, yet existing benchmarks often overlook the complex physical reasoning required for point-and-click puzzle games. This paper introduces Vision-Language Against The Incredible Machine (VLATIM), a benchmark designed to evaluate human-like logical problem-solving capabilities within the classic physics puzzle game The Incredible Machine 2 (TIM). Unlike existing benchmarks, VLATIM specifically targets the critical gap between high-level logical reasoning and continuous action spaces requiring precise mouse interactions. This benchmark is structured into five progressive parts, assessing capabilities that range from basic visual grounding and domain understanding to multi-step manipulation and full puzzle solving. Our results reveal a significant disparity between reasoning and execution. While large proprietary models demonstrate superior planning abilities, they struggle with precise visual grounding. Consequently, they do not yet show human-like problem-solving capabilities.

视觉语言模型物理推理人机交互

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。