让视觉与文本在隐空间动态交织,提升多模态推理能力
Reasoning Within the Mind: Dynamic Multimodal Interleaving in Latent Space
- 通过置信度引导的隐空间策略优化,动态调整思维令牌
- 在七项基准测试中超越现有方法,推理准确率提升显著
- 适合需要高效高精度多模态推理的研究者与开发者
近期多模态大语言模型(MLLM)通过在语义空间中引入思维链(CoT)推理,显著提升了跨模态理解与推理能力。部分研究将CoT机制扩展至视觉模态,通过外部工具或显式图像生成来整合视觉信息,但这些方法仍依赖于显式的逐步推理,存在感知-推理交互不稳定、计算开销大等问题。受人类认知启发,我们提出一种测试时动态多模态隐空间推理框架(DMLR),其核心为置信度引导的隐空间策略梯度优化,用于精细化调整隐空间思维令牌。同时引入动态视觉注入策略,在每一步隐空间思维令牌生成时,检索最相关的视觉特征,并更新最佳视觉块集合,将其注入思维令牌以实现动态视觉-文本交织。在七个多模态推理基准及多种模型架构上的实验表明,DMLR显著提升推理与感知性能,同时保持高推理效率。
原文摘要 · Abstract (English)
Recent advancements in Multimodal Large Language Models (MLLMs) have significantly enhanced cross-modal understanding and reasoning by incorporating Chain-of-Thought (CoT) reasoning in the semantic space. Building upon this, recent studies extend the CoT mechanism to the visual modality, enabling models to integrate visual information during reasoning through external tools or explicit image generation. However, these methods remain dependent on explicit step-by-step reasoning, unstable perception-reasoning interaction and notable computational overhead. Inspired by human cognition, we posit that thinking unfolds not linearly but through the dynamic interleaving of reasoning and perception within the mind. Motivated by this perspective, we propose DMLR, a test-time Dynamic Multimodal Latent Reasoning framework that employs confidence-guided latent policy gradient optimization to refine latent think tokens for in-depth reasoning. Furthermore, a Dynamic Visual Injection Strategy is introduced, which retrieves the most relevant visual features at each latent think token and updates the set of best visual patches. The updated patches are then injected into latent think token to achieve dynamic visual-textual interleaving. Experiments across seven multimodal reasoning benchmarks and various model architectures demonstrate that DMLR significantly improves reasoning and perception performance while maintaining high inference efficiency.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。