arXiv:2606.19965cs.CVcs.AI2026-06

测试多模态模型从视觉信息到动作的转换能力,发现其在不同任务下表现差异巨大。

ROSE: Benchmarking the Perception-to-Action Gap in Multimodal Models

论文配图:ROSE: Benchmarking the Perception-to-Action Gap in Multimodal Models
图 1 · 摘自论文原文
  • 固定视觉场景,通过变化区域约束和符号输出测试模型推理能力。
  • 模型在动作任务上性能下降44.5个百分点,人类仍保持98.8%准确率。
  • 揭示模型在共享视觉信息转为上下文相关动作时存在固有瓶颈。

多模态大语言模型日益被期望基于视觉信息采取行动,但同一场景在不同任务背景下可能需要不同动作。模型能否可靠地将相同视觉证据转化为当前上下文所需的动作?为此,我们提出 extsc{ROSE}(参考条件下的奇异性与符号执行),一个控制性基准,固定视觉场景,同时变化区域约束和所需的符号输出。通过耦合计数与坐标动作任务, extsc{ROSE} 测试模型在不同上下文中是否能推断隐含的多数参考,并基于细粒度视觉证据采取行动。在九个近期的多模态大模型中,性能从以计数为导向的任务下降至区域条件动作任务达44.5个百分点,而人类表现仍高达98.8%。该差距在成对场景和区域上依然存在,即使同一模型对计数任务给出正确答案;全局点击与匹配局部对照实验表明,坐标定位仅解释部分性能损失,揭示出一种独立于模型、且依赖于模型自身的瓶颈——即将共享视觉证据转化为特定上下文动作的能力不足。

原文摘要 · Abstract (English)

Multimodal large language models (MLLMs) are increasingly expected to act on visual information, yet the same scene may require different actions under different task contexts. How reliably can a model turn the same visual evidence into the action required by the current context? To answer this question, we introduce \textsc{ROSE} (\textbf{R}eference-conditioned \textbf{O}ddity and \textbf{S}ymbolic \textbf{E}xecution), a controlled benchmark that holds the visual scene fixed while varying region constraints and required symbolic outputs. Through coupled counting and coordinate-action tasks, \textsc{ROSE} tests whether models can infer an implicit majority reference and act on the resulting fine-grained visual evidence under changing contexts. Across nine recent MLLMs, performance drops by as much as 44.5 percentage points from counting-oriented tasks to region-conditioned action, despite 98.8\% human performance. The gap persists on paired scenes and regions for which the same model returns the correct count, while global-click and matched local controls show that coordinate grounding explains only part of the loss, revealing a distinct, model-dependent bottleneck in turning shared visual evidence into context-specific actions.

多模态模型评估动作推理

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。