让被压制的视觉隐变量重获推理能力,不改模型直接提升多模态理解。
Visual Latents Know More Than They Say: Unsilencing Latent Reasoning in MLLMs

- 训练中视觉隐变量被抑制,推理时通过优化释放其潜力。
- 在8个基准上,零参数更新下准确率显著提升。
- 适合关注模型内在推理机制与可解释性的研究者。
连续隐空间推理为多模态模型提供了紧凑的替代方案,避免使用显式思维链令牌,能整合高维视觉证据。然而我们发现现有视觉隐变量推理方法存在未被察觉的优化病理:尽管视觉隐变量在训练中语义逐渐丰富,但其对最终答案预测的贡献却被系统性抑制。在共享参数空间中,自回归目标倾向于依赖直接视觉输入,使隐变量趋向过渡态而非信息丰富的推理内容。我们称此现象为‘沉默的视觉隐变量’。为解决该问题,我们在推理时解耦两个冲突目标,直接优化隐变量推理,保持主干参数冻结。第一阶段通过查询引导的对比隐变量-视觉对齐,提升隐变量语义质量并防止隐变量坍塌;第二阶段通过置信度进展奖励,激励隐变量段落上的预测分布逐步集中,引导预测经由隐变量推理而非绕过。在八个基准和四个模型主干上的实验表明,无需任何参数更新,推理时的隐变量优化可有效释放被抑制的推理能力。
原文摘要 · Abstract (English)
Continuous latent-space reasoning offers a compact alternative to textual chain-of-thought for multimodal models, enabling high-dimensional visual evidence to be integrated without explicit reasoning tokens. However, we identify a previously overlooked optimization pathology in existing latent visual reasoning methods: although visual latents become semantically enriched during training, their contribution to final answer prediction is systematically suppressed. Within the shared parameter space, the autoregressive objective favors shortcut reliance on direct visual input, driving latent tokens toward transition-like states rather than informative reasoning content. We term this phenomenon Silenced Visual Latents. To address it, we disentangle the two conflicting objectives by directly optimizing the latent reasoning at inference time, keeping backbone parameters frozen. In Stage I, visual latents are warmed up via query-guided contrastive latent--visual alignment, improving semantic quality while preventing latent collapse. In Stage II, the latent reasoning is further optimized via a confidence-progression reward, which incentivizes predicted token distributions along the latent span to become progressively more concentrated, routing predictions through the latent reasoning rather than bypassing it. Experiments across eight benchmarks and four model backbones show that inference-time latent optimization, without any parameter updates, effectively unleashes the suppressed reasoning capacity of visual latents.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。