arXiv:2604.07518cs.CL2026-04中稿 · EMNLP

让视觉语言模型分步看图推理,提升复杂问题理解能力。

Decompose, Look, and Reason: Reinforced Latent Reasoning for VLMs

  • 动态拆解问题,提取条件相关的视觉隐变量
  • 在多个基准上超越现有方法,推理过程更可解释
  • 适合需要精细视觉推理的应用场景

视觉语言模型在复杂视觉推理任务中常因文本思维链(CoT)导致视觉信息丢失。现有方法或增加工具调用开销,或依赖局部化图像块嵌入,难以支持多步推理中的语义提取。本文提出「分解、查看、推理」(DLR)框架,通过三阶段训练和新颖的球形高斯隐变量策略,动态将查询分解为文本前提,生成前提相关的连续视觉隐变量,并基于有依据的推理链得出答案。大量实验表明,DLR在以视觉为中心的基准上持续优于强基线,包括纯文本、交错多模态CoT及隐变量推理方法,同时具备更优的逐步可解释性。

原文摘要 · Abstract (English)

Vision-Language Models often struggle with complex visual reasoning due to the visual information loss in textual CoT. Existing methods either add the cost of tool calls or rely on localized patch-based embeddings that are insufficient to extract semantics in multi-step reasoning. We propose "Decompose, Look, and Reason" (DLR), a reinforced latent reasoning framework that dynamically decomposes queries into textual premises, extracts premise-conditioned continuous visual latents, and deduces answers through grounded rationales. We introduce a three-stage training pipeline and propose a novel Spherical Gaussian Latent Policy, to enable effective exploration in the latent space. Extensive experiments on vision-centric benchmarks show that DLR consistently outperforms strong baselines, including text-only, interleaved multimodal CoT, and latent reasoning methods, while providing superior stepwise interpretability.

视觉推理多模态隐变量可解释性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。