arXiv:2602.05359cs.CV2026-02被引 1

让AI在隐空间中分步推理,用视觉线索逐步提升理解力。

Multimodal Latent Reasoning via Hierarchical Visual Cues Injection

  • 通过递归扩展Transformer,在隐空间内实现多步迭代推理。
  • 注入从全局到局部的视觉线索,显著提升复杂场景理解能力。
  • 适合需要精准视觉推理的场景,如医疗影像分析或自动驾驶。

多模态大语言模型虽具备出色感知能力,但其推理常依赖端到端生成或以语言为中心的思维链(CoT),存在效率低、冗长且易幻觉的问题。本文提出基于层次化视觉线索注入的多模态隐空间推理框架(HIVE),在不依赖表面文本推理的前提下,引入“慢思考”机制。方法通过递归扩展Transformer块,构建内部迭代回路,将全局场景上下文至细粒度区域细节的视觉线索直接注入模型隐表示中,实现完全在对齐隐空间内的有根据多步推理。大量实验证明,测试时引入视觉知识可有效提升性能,融合层次化信息显著增强模型对复杂场景的理解能力。

原文摘要 · Abstract (English)

The advancement of multimodal large language models (MLLMs) has enabled impressive perception capabilities. However, their reasoning process often remains a "fast thinking" paradigm, reliant on end-to-end generation or explicit, language-centric chains of thought (CoT), which can be inefficient, verbose, and prone to hallucination. This work posits that robust reasoning should evolve within a latent space, integrating multimodal signals seamlessly. We propose multimodal latent reasoning via HIerarchical Visual cuEs injection (\emph{HIVE}), a novel framework that instills deliberate, "slow thinking" without depending on superficial textual rationales. Our method recursively extends transformer blocks, creating an internal loop for iterative reasoning refinement. Crucially, it injectively grounds this process with hierarchical visual cues from global scene context to fine-grained regional details directly into the model's latent representations. This enables the model to perform grounded, multi-step inference entirely in the aligned latent space. Extensive evaluations demonstrate that test-time scaling is effective when incorporating vision knowledge, and that integrating hierarchical information significantly enhances the model's understanding of complex scenes.

多模态推理隐空间视觉线索深度模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。