arXiv:2511.02360cs.CVcs.CL2025-11被引 2

让视觉聚焦在隐空间完成,提速60%还更准

LaRe: Latent Refocusing for Multimodal Reasoning

  • 在隐空间内实现视觉焦点动态调整,避免显式图像裁剪
  • 相比基线提升7.6%准确率,推理token减少59.7%
  • 适合追求高效多模态推理的模型部署者

思维链(CoT)推理通过分解复杂任务提升逻辑性能,但其多模态扩展面临权衡。现有“看图思考”范式通过显式裁剪图像区域实现视觉聚焦,但计算开销迅速增加。新兴的隐空间推理虽降低令牌消耗,却缺乏动态聚焦能力。本文认为该权衡源于默认假设:有效视觉聚焦必须以显式令牌形式实现。基于此,提出隐空间聚焦(LaRe),使视觉聚焦完全在隐空间中完成。进一步设计语义增强训练策略,通过视觉重建目标保持隐空间语义结构。实验表明,LaRe相比现有基线平均准确率提升7.6%,推理所需令牌数减少59.7%。当扩展至80亿参数视觉-语言模型时,性能达到当前最优水平,验证了该隐空间聚焦范式的有效性。

原文摘要 · Abstract (English)

Chain of Thought (CoT) reasoning enhances logical performance by decomposing complex tasks, yet its multimodal extension faces a trade-off. The prevailing Thinking with Images paradigm achieves visual refocusing by explicitly cropping image regions, yet incurs rapidly growing computational overhead. The emerging line of latent-space reasoning reduces token consumption, but lacks the capacity for dynamic refocusing. We argue that this trade-off stems from a tacitly accepted premise that effective visual refocusing must occur in the form of explicit tokens. Building on this, we propose Latent Refocusing (LaRe), a new multimodal reasoning paradigm in which visual refocusing takes place entirely within the latent space. We further design a semantic augmentation training strategy that ensures the semantic structure of the latent space through visual reconstruction objective. Experimental evaluations demonstrate that LaRe improves average accuracy by 7.6% compared to existing baselines while reducing the number of tokens required for inference by 59.7%. When scaled to a 8B-parameter Vision-Language Model backbone, LaRe achieves performance comparable to state-of-the-art methods, demonstrating the efficacy of our proposed latent refocusing paradigm for multimodal reasoning.

多模态推理隐空间视觉聚焦

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。