arXiv:2605.07106cs.CL2026-05被引 2

让大模型在视觉推理中更精准地理解空间与语义信息。

Retrieve, Integrate, and Synthesize: Spatial-Semantic Grounded Latent Visual Reasoning

论文配图:Retrieve, Integrate, and Synthesize: Spatial-Semantic Grounded Latent Visual Reasoning
图 1 · 摘自论文原文
  • 通过空间-语义双重锚定,让隐变量推理与模型原有逻辑兼容。
  • 在多个数据集上超越主流基线,尤其在细粒度推理任务中提升显著。
  • 适合追求可解释性与高精度视觉推理的开发者与研究者。

多模态大语言模型在视觉语言推理方面取得显著进展,但多数方法将视觉证据压缩为离散文本思想,造成细粒度感知的信息瓶颈。近期的隐变量推理方法尝试在连续隐藏状态中推理,但我们发现它们存在流形不兼容问题:隐变量轨迹偏离预训练推理路径,趋于泛化模式,且常被答案生成阶段跳过。为此,我们提出RIS(Retrieve, Integrate, and Synthesize)框架,构建空间-语义双重锚定的隐变量推理机制,作为预训练MLLM计算的兼容扩展。我们首先构建一个带边界框和区域语义描述的分步推理数据集。基于此监督信号,RIS将隐变量锚定于空间与语义证据,通过渐进注意力瓶颈强化其因果作用,并引入短语言过渡令牌,将合成的隐状态回传至词汇对齐解码。在V*、HRBench4K、HRBench8K、MMVP和BLINK上的实验表明,RIS持续优于闭源/开源及隐变量推理基线。进一步分析显示,RIS学习到多样化、可解释且逐步融合的隐变量轨迹,为实现可信的MLLM内部视觉推理提供了可行路径。

原文摘要 · Abstract (English)

Multimodal Large Language Models (MLLMs) have made remarkable progress on vision-language reasoning, yet most methods still compress visual evidence into discrete textual thoughts, creating an information bottleneck for fine-grained perception. Recent latent visual reasoning methods attempt to reason in continuous hidden states, but we find that they suffer from insufficient manifold compatibility: latent trajectories drift away from pretrained reasoning circuits, collapse into instance-agnostic patterns, and are often bypassed during answer generation. To address these issues, we propose RIS (Retrieve, Integrate, and Synthesize), a spatial-semantic grounded framework that develops latent reasoning as a compatible extension of pretrained MLLM computation. We first construct a step-wise grounded reasoning dataset with bounding boxes and region-specific semantic descriptions. Built on this supervision, RIS anchors latent tokens to both spatial and semantic evidence, enforces their causal role through a progressive attention bottleneck, and introduces short language transition tokens to bridge synthesized latent states back to vocabulary-aligned decoding. Experiments on V*, HRBench4K, HRBench8K, MMVP, and BLINK show consistent improvements over closed/open-source and latent reasoning baselines. Further analyses demonstrate that RIS learns diverse, interpretable, and progressively integrated latent trajectories, offering a practical path toward faithful internal visual reasoning in MLLMs.

视觉推理隐变量多模态

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。