arXiv:2606.01287cs.CVcs.AI2026-06被引 1

揭示视觉推理中潜伏标记的真实作用,打破视觉记忆的误解。

Beyond Visual Memory: Mechanistic Diagnostics of Latent Visual Reasoning

论文配图:Beyond Visual Memory: Mechanistic Diagnostics of Latent Visual Reasoning
图 1 · 摘自论文原文
  • 分解潜伏标记为槽位、边界标记和格式三部分进行验证
  • 仅保留边界标记即可维持78%至100%的性能提升
  • 适合关注模型内在机制与训练监督影响的研究者

近期的潜伏视觉推理方法通过在多模态语言模型中插入连续潜伏标记取得了显著进步。这些进步通常归因于标记编码视觉证据;然而,最新分析揭示了一个悖论:这些标记与图像关联松散,对答案贡献甚微。关键在于,现有分析将潜伏标记视为单一整体,掩盖了增益来源。为此,我们将其分解为三个可检验组件:潜伏槽位、边界标记和格式,并构建了在理想条件下表现最优的探测方法。在六种方法阶段设置和四个感知密集型基准上,潜伏槽位完全无法支持视觉记忆假设。令人惊讶的是,仅保留边界标记即可在多个设置中保持78%至100%的性能增益,且模型在潜伏位置的注意力范围比答案位置更窄。结果表明,槽位内容并非可恢复的视觉记忆;大部分收益实则源于标记与格式控制,以及视觉注意力路由。在匹配准确率下,不同方法仍依赖显著不同的机制,由训练监督塑造。因此,潜伏视觉推理不仅需以准确率评估,更应考察模型实际依赖的机制。

原文摘要 · Abstract (English)

Recent latent visual reasoning methods achieve substantial gains by inserting continuous latent tokens into multimodal language models. These gains are commonly attributed to the tokens encoding visual evidence; recent analyses, however, reveal a paradox: the tokens are loosely tied to the image and contribute little to the answer. Critically, these analyses treat latent tokens as a single unit, obscuring the source of the gains. We therefore decompose latent tokens into three testable components: latent slots, boundary markers, and format, and develop a state-of-the-art method as a probe under favorable conditions. Across six method-stage settings and four perception-heavy benchmarks, latent slots fail every prediction of the visual-memory account. Strikingly, retaining only the boundary markers preserves 78 to 100% of the gain in several settings, while the model attends to the image more narrowly at latent positions than at answer positions. These results do not support slot contents as recoverable visual memory; much of the benefit is instead associated with marker and format control, together with visual-attention routing. At matched accuracy, methods can still rely on markedly different mechanisms shaped by training supervision. Latent visual reasoning thus needs evaluation not only by accuracy but by what the model actually relies on.

视觉推理机制诊断注意力路由模型解释

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。