arXiv:2511.17106cs.CV2025-11被引 3

用动态视觉提示压缩多模态推理链,更短更准

ChainV: Atomic Visual Hints Make Multimodal Reasoning Shorter and Better

  • 根据注意力强度选关键视觉片段,动态融入推理过程
  • 数学类任务准确率提升2.3%,推理延迟降低51.4%
  • 适合需要视觉辅助的多步符号推理场景

近期多模态推理模型在文本与视觉任务中表现优异,但顶尖模型在生成长推理链时仍存在冗余自我反思问题。尽管训练无关的思维链压缩方法已在大语言模型中出现,但其依赖静态视觉参考,对多模态推理增益有限。为此,我们提出ChainV框架,通过动态整合视觉提示优化推理过程,使多模态推理更短更准。ChainV首先基于前一推理步骤粗选视觉区块,再依据平均注意力强度筛选最具代表性的原子级视觉提示;同时引入一致性评估机制判断提示可靠性,引导模型自适应调整反思程度。最终,选定视觉提示的像素坐标及其可靠性通过伯努利随机过程融入思考。实验表明,该方法显著提升推理准确率与效率,尤其在依赖视觉的数学密集型任务中表现突出。例如,在MIMO-VL-RL上的MathVista数据集上,准确率提升2.3%,推理延迟下降51.4%,输出词元长度减少24.5%。

原文摘要 · Abstract (English)

Recent advances in multimodal reasoning models have demonstrated impressive capabilities across text and vision. However, even leading models exhibit redundant self-reflection when generating lengthy reasoning chains. While training-free CoT compression methods have emerged in the LLMs domain, they rely on static visual references and thus provide limited gains for multimodal reasoning. Therefore, we propose ChainV, a framework that dynamically integrates visual hints into the reasoning process, thereby making multimodal reasoning shorter and better. Specifically, ChainV first performs a coarse visual patch selection based on the previous reasoning step, then refines it by identifying the most representative atomic visual hint according to the averaged attention intensity. Additionally, ChainV introduces a consistency-based evaluation mechanism to assess the reliability of the chosen hint, guiding the model to adaptively adjust its level of self-reflection. Eventually, the pixel coordinates of the selected visual hint and its reliability are incorporated into thinking with a Bernoulli stochastic process. Experiments indicate that our method significantly improves reasoning accuracy and efficiency, especially on math-intensive benchmarks where visual hints are crucial for multi-step symbolic reasoning. For example, ChainV achieves $2.3\%$ improvement on the MathVista within MIMO-VL-RL, while reducing inference latency by $51.4\%$ and shortening output token length by $24.5\%$.

多模态推理视觉提示思维链压缩数学推理

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。