检验多模态模型是否真用中间视觉状态进行推理
See2Think: Do Multimodal Models Really Use Intermediate Visual States?

- 构建统一评估框架,记录模型推理中的视觉操作与状态变化
- 12类任务中,视觉状态依赖导致准确率下降超10个百分点
- 模型能选对操作但渲染不实,反馈多未必更准
多模态大语言模型在推理中越来越多地使用草图、标注、工具和中间图像,但其是否真正依赖这些视觉状态仍不明确。现有基准存在任务覆盖窄、部分问题可纯文本求解,以及评价仅关注最终答案而忽略中间视觉状态生成、渲染与使用的问题。我们提出See2Think,包含See2ThinkBench和视觉思维过程(VAoT)的统一评估框架。See2ThinkBench包含1,200个开放性、视觉依赖型问题,覆盖12个任务类别,涵盖二维结构、三维场景与真实世界推理。VAoT在四种受控推理设置下记录文本思考、视觉操作、渲染状态及后续推理。评估主流专有与开源多模态模型发现,视觉推理强依赖模型与环境,无单一设置在所有任务中占优。过程分析显示,模型通常能选择相关视觉操作,但忠实渲染仍是主要瓶颈;高反馈采纳率并不必然带来准确率提升。在任务相关的损坏反馈下,模型行为明显依赖视觉状态,准确率下降超过10个百分点。
原文摘要 · Abstract (English)
Multimodal large language models increasingly use sketches, annotations, tools, and intermediate images during reasoning, but it remains unclear whether they truly rely on these visual states. Existing benchmarks are limited both by task collections with narrow coverage or partially text-solvable samples and by evaluations that emphasize final answers without diagnosing how intermediate visual states are generated, rendered, and used. We introduce See2Think, a unified evaluation framework comprising See2ThinkBench and Visual Action-of-Thought (VAoT). See2ThinkBench contains 1,200 open-ended, visually dependent problems across 12 task categories spanning 2D structured, 3D scene, and real-world reasoning. VAoT records textual thoughts, visual actions, rendered states, and subsequent reasoning under four controlled inference settings. Evaluating representative proprietary and open-source multimodal models, we find that visual reasoning is strongly model- and environment-dependent, with no single setting consistently dominating across tasks. Process analysis further shows that models usually select relevant visual operations, while faithful rendering remains the clearest bottleneck and high feedback uptake does not necessarily translate into accuracy gains. Under task-relevant corrupted feedback, models exhibit behavioral dependence on visual states, with accuracy dropping by over 10 percentage points in controlled interventions.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。