arXiv:2602.02465cs.AIcs.CV2026-02中稿 · ICML被引 6

测试大模型用视觉想象推理的能力,发现效果不佳。

MentisOculi: Revealing the Limits of Reasoning with Mental Imagery

  • 设计分层多步推理任务,检验模型生成和使用视觉图像的能力。
  • 无论用隐式特征还是显式图像,模型性能均未提升。
  • 即使给正确图像,模型仍因生成错误而无法利用,适合研究推理局限性的人看。

前沿模型正从仅处理视觉信息的多模态大语言模型(MLLMs)转向能原生交替生成的统一多模态模型(UMMs)。这一转变激发了将中间可视化作为推理辅助工具的兴趣,类似于人类的心理意象。核心在于能否以目标导向的方式形成、维持和操作视觉表征。为此,我们开发了MentisOculi——一个程序化、分层的多步推理任务集,可被视觉方式解决,专为挑战前沿模型而设计。评估从潜在令牌到显式生成图像的各种视觉策略,发现它们普遍未能提升性能。对UMMs的分析揭示了一个关键局限:尽管具备解决任务的文本推理能力,且有时能生成正确图像,但模型会累积生成错误,甚至无法利用真实图像。研究结果表明,尽管视觉思维具有内在吸引力,但尚未真正提升模型推理。MentisOculi为跨多种模型家族分析并弥合这一差距奠定了必要基础。

原文摘要 · Abstract (English)

Frontier models are transitioning from multimodal large language models (MLLMs) that merely ingest visual information to unified multimodal models (UMMs) capable of native interleaved generation. This shift has sparked interest in using intermediate visualizations as a reasoning aid, akin to human mental imagery. Central to this idea is the ability to form, maintain, and manipulate visual representations in a goal-oriented manner. To evaluate and probe this capability, we develop MentisOculi, a procedural, stratified suite of multi-step reasoning problems amenable to visual solution, tuned to challenge frontier models. Evaluating visual strategies ranging from latent tokens to explicit generated imagery, we find they generally fail to improve performance. Analysis of UMMs specifically exposes a critical limitation: While they possess the textual reasoning capacity to solve a task and can sometimes generate correct visuals, they suffer from compounding generation errors and fail to leverage even ground-truth visualizations. Our findings suggest that despite their inherent appeal, visual thoughts do not yet benefit model reasoning. MentisOculi establishes the necessary foundation to analyze and close this gap across diverse model families.

多模态模型推理能力视觉想象

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。