arXiv:2602.22766cs.CL2026-02被引 6

发现隐空间推理无效,提出用文本显式想象更有效

Imagination Helps Visual Reasoning, But Not Yet in Latent Space

  • 用因果中介分析揭示隐状态对输入不敏感、对答案影响小
  • 隐状态编码视觉信息少且高度相似,无法支撑推理
  • 新方法CapImagine通过文本显式想象,性能显著更优

隐空间视觉推理旨在通过多模态大模型的隐藏状态模拟人类想象力。尽管该范式被认为有前景,其有效性来源仍不明确。为探究其真实机制,我们采用因果中介分析,将输入视为处理变量,隐状态为中介变量,最终答案为结果变量。研究发现两个关键断链:(a) 输入-隐状态断链:对输入进行剧烈扰动后,隐状态变化极小,表明隐状态未有效关注输入序列;(b) 隐状态-答案断链:对隐状态扰动几乎不影响最终答案,说明其对结果的因果影响微弱。进一步探查显示,隐状态编码的视觉信息有限且高度相似。因此我们质疑隐空间推理的必要性,提出名为CapImagine的新方法,引导模型通过文本显式生成想象。在多个视觉基准上实验表明,CapImagine显著优于复杂的隐空间基线,凸显显式想象在视觉推理中的更强潜力。

原文摘要 · Abstract (English)

Latent visual reasoning aims to mimic human's imagination process by meditating through hidden states of Multimodal Large Language Models. While recognized as a promising paradigm for visual reasoning, the underlying mechanisms driving its effectiveness remain unclear. Motivated to demystify the true source of its efficacy, we investigate the validity of latent reasoning using Causal Mediation Analysis. We model the process as a causal chain: the input as the treatment, the latent tokens as the mediator, and the final answer as the outcome. Our findings uncover two critical disconnections: (a) Input-Latent Disconnect: dramatic perturbations on the input result in negligible changes to the latent tokens, suggesting that latent tokens do not effectively attend to the input sequence. (b) Latent-Answer Disconnect: perturbations on the latent tokens yield minimal impact on the final answer, indicating the limited causal effect latent tokens imposing on the outcome. Furthermore, extensive probing analysis reveals that latent tokens encode limited visual information and exhibit high similarity. Consequently, we challenge the necessity of latent reasoning and propose a straightforward alternative named CapImagine, which teaches the model to explicitly imagine using text. Experiments on vision-centric benchmarks show that CapImagine significantly outperforms complex latent-space baselines, highlighting the superior potential of visual reasoning through explicit imagination.

视觉推理隐空间显式想象大模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。