arXiv:2604.10500cs.CV2026-04被引 2

让视觉信息在多模态推理中不被忽视,提升复杂问题的推理深度与速度。

Visual Enhanced Depth Scaling for Multimodal Latent Reasoning

论文配图:Visual Enhanced Depth Scaling for Multimodal Latent Reasoning
图 1 · 摘自论文原文
  • 用视觉重放和动态深度扩展,解决视觉特征优化不足的问题。
  • 在多个基准上达到顶尖性能,推理速度比传统方法快3倍以上。
  • 适合需要高效高精度多模态推理的场景,如视觉问答与图像理解。

多模态潜在推理作为一种新兴范式,以隐式特征传播替代显式的思维链(CoT)解码,同时提升表征丰富性并降低推理延迟。通过分析潜在训练中的词元级梯度动态,我们发现两个关键现象:(1) 由于固有的语言偏见,视觉词元的梯度范数显著小于文本词元,导致系统性视觉欠优化;(2) 语义简单的词元快速收敛,而复杂词元表现出持续的梯度不稳定性,受限于固定架构深度。为应对这些局限,我们提出视觉重放模块与路由深度缩放机制,协同增强视觉感知并精细化复杂潜在表示以支持更深层上下文推理。前者利用因果自注意力估计词元显著性,通过空间一致约束强化细粒度定位;后者自适应为复杂词元分配额外推理步骤,实现更深上下文优化。在逐步内化显式CoT到紧凑潜在表示的课程学习策略引导下,本框架在多个基准上达到最先进性能,并相比显式CoT基线实现显著推理加速。

原文摘要 · Abstract (English)

Multimodal latent reasoning has emerged as a promising paradigm that replaces explicit Chain-of-Thought (CoT) decoding with implicit feature propagation, simultaneously enhancing representation informativeness and reducing inference latency. By analyzing token-level gradient dynamics during latent training, we reveal two critical observations: (1) visual tokens exhibit significantly smaller gradient norms than their textual counterparts due to inherent language bias, resulting in systematic visual under-optimization; and (2) semantically simple tokens converge rapidly, whereas complex tokens exhibit persistent gradient instability constrained by fixed architectural depths. To address these limitations, we propose a visual replay module and routing depth scaling to collaboratively enhance visual perception and refine complicated latents for deeper contextual reasoning. The former module leverages causal self-attention to estimate token saliency, reinforcing fine-grained grounding through spatially-coherent constraints. Complementarily, the latter mechanism adaptively allocates additional reasoning steps to complex tokens, enabling deeper contextual refinement. Guided by a curriculum strategy that progressively internalizes explicit CoT into compact latent representations, our framework achieves state-of-the-art performance across diverse benchmarks while delivering substantial inference speedups over explicit CoT baselines.

多模态推理视觉增强深度缩放潜空间

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。