arXiv:2605.11856cs.CVcs.CL2026-05被引 4

让视觉推理像看图思考一样统一高效,减少冗余文本

UniVLR: Unifying Text and Vision in Visual Latent Reasoning for Multimodal LLMs

论文配图:UniVLR: Unifying Text and Vision in Visual Latent Reasoning for Multimodal LLMs
图 1 · 摘自论文原文
  • 将文本推理与视觉证据融合为共享视觉工作区
  • 推理过程仅通过视觉隐变量完成,生成的推理令牌减少60%以上
  • 适合需要高效视觉推理的多模态大模型应用

多模态大语言模型越来越需要具备图像思维能力,但现有视觉隐变量推理方法仍依赖在视觉隐变量中穿插显式文本思维链。这种交替设计限制了效率,且使推理在独立的文本与视觉通道间碎片化。我们提出UniVLR,一种统一的视觉隐变量推理框架,将文本推理与辅助视觉证据视为共享的视觉工作区。不再保留文本思维链作为独立的推理路径,UniVLR将推理轨迹与辅助图像一起渲染,并学习将其压缩为紧凑的视觉隐变量令牌。推理时,模型仅通过视觉隐变量进行推断,并直接解码最终答案,避免外部工具调用和冗长的文本推理。在真实世界感知与视觉推理任务上的实验表明,UniVLR优于以往视觉隐变量推理方法,同时生成的推理令牌显著减少,表明其为多模态大模型视觉思维提供了一种更统一、高效的范式。

原文摘要 · Abstract (English)

Multimodal large language models are increasingly expected to perform thinking with images, yet existing visual latent reasoning methods still rely on explicit textual chain-of-thought interleaved with visual latent tokens. This interleaved design limits efficiency and keeps reasoning fragmented across separate text and vision channels. We propose UniVLR, a unified visual latent reasoning framework that treats textual reasoning and auxiliary visual evidence as a shared visual workspace. Instead of preserving text CoT as an independent inference-time path, UniVLR renders reasoning traces together with auxiliary images and learns to compress this unified representation into compact visual latent tokens. At inference time, the model reasons only through visual latents and directly decodes the final answer, avoiding both external tool calls and verbose text reasoning. Experiments on real-world perception and visual reasoning tasks show that UniVLR outperforms prior visual latent reasoning methods while using substantially fewer generated reasoning tokens, suggesting a more unified and efficient paradigm for visual thinking in MLLMs.

多模态视觉推理隐变量大模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。