arXiv:2510.12603cs.CVcs.AI2025-10ACL被引 20

在隐空间中融合视觉与文本推理,提升多模态大模型效率与准确率。

Reasoning in the Dark: Interleaved Vision-Text Reasoning in Latent Space

  • 在隐空间中交替注入视觉与文本信息,实现无需标注的推理过程。
  • 在M³CoT和ScienceQA上准确率提升5.45%,推理速度加快5倍以上。
  • 适合追求高效多模态推理的开发者与研究者使用。

多模态推理通过在最终答案前引入中间推理步骤,增强多模态大语言模型(MLLMs)的能力。其发展已从纯文本推理扩展至融合视觉信息,使思维过程可通过图像与文本共同表达。然而,现有方法依赖人工标注的显式推理步骤,导致标注成本高且推理延迟显著。为此,我们提出多模态隐空间推理,兼具多模态表征优势、减少标注需求与高效推理。具体地,我们设计了交错视觉-文本隐空间推理(IVT-LR),在隐空间中穿插注入视觉与文本信息。每个推理步骤由两部分隐式构成:前一步的隐式文本(隐藏状态)与一组精选的图像嵌入(隐式视觉)。我们进一步提出渐进式多阶段训练策略,使MLLMs能执行上述多模态隐空间推理。在M³CoT与ScienceQA数据集上的实验表明,所提方法平均准确率提升5.45%,同时推理速度相较现有方法提升超过5倍。

原文摘要 · Abstract (English)

Multimodal reasoning aims to enhance the capabilities of MLLMs by incorporating intermediate reasoning steps before reaching the final answer. It has evolved from text-only reasoning to the integration of visual information, enabling the thought process to be conveyed through both images and text. Despite its effectiveness, current multimodal reasoning methods depend on explicit reasoning steps that require labor-intensive vision-text annotations and inherently introduce significant inference latency. To address these issues, we introduce multimodal latent reasoning with the advantages of multimodal representation, reduced annotation, and inference efficiency. To facilitate it, we propose Interleaved Vision-Text Latent Reasoning (IVT-LR), which injects both visual and textual information in the reasoning process within the latent space. Specifically, IVT-LR represents each reasoning step by combining two implicit parts: latent text (the hidden states from the previous step) and latent vision (a set of selected image embeddings). We further introduce a progressive multi-stage training strategy to enable MLLMs to perform the above multimodal latent reasoning steps. Experiments on M$^3$CoT and ScienceQA demonstrate that our IVT-LR method achieves an average performance increase of 5.45\% in accuracy, while simultaneously achieving a speed increase of over 5 times compared to existing approaches.

多模态推理隐空间效率提升视觉语言

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。