通过信息论分析优化视觉信息传播,提升长链多模态推理的准确性。
Reflection Anchors for Propagation-Aware Visual Retention in Long-Chain Multimodal Reasoning

- 基于信息论推导干预收益下界,指导选择高熵视觉锚点。
- 在多个基准上显著超越强基线,提升长链推理中视觉依赖性。
- 适合需要精准视觉记忆的复杂多模态任务研究者。
长链思维(CoT)推理虽能提升大视觉-语言模型性能,但视觉信息在生成过程中常会衰减,限制了长时序多模态推理。现有方法要么在推理时重注视觉信息,要么训练强化对齐的策略,但干预位置依赖感知启发式,且局部视觉影响如何传播缺乏明确建模。本文从信息论角度分析,推导出单步干预对下游视觉增益的下界,揭示两个关键因素:局部分支空间(令牌熵)与下游视觉传播潜力(后缀相对于视觉边缘化参考的差异)。基于此,提出反射锚点策略优化(RAPO),一种基于GRPO的策略优化方法,选择高熵反射锚点,并优化链路掩码的有限窗口KL代理以增强下游视觉依赖。在推理密集型及通用领域基准上的实验表明,RAPO在多个LVLM主干网络上均显著优于强基线。机制分析进一步显示,反射锚点集中于视觉敏感决策点,且RAPO增强了生成轨迹中的对比性视觉依赖信号。
原文摘要 · Abstract (English)
Long chain-of-thought (CoT) reasoning improves large vision--language models, but visual information often fades during generation, limiting long-horizon multimodal reasoning. Existing methods either re-inject vision at inference or train policies for stronger grounding, but where to intervene relies on perception heuristics rather than principled gain analysis, and how local visual influence propagates remains implicit. We study this problem from an information-theoretic standpoint and derive a lower bound on the downstream visual gain of a one-step intervention, which suggests two factors: local branching room (token entropy) and downstream visual propagation potential (suffix divergence from a vision-marginalized reference). Guided by this analysis, we propose reflection-anchor policy optimization (RAPO), a GRPO-based policy optimization method that selects high-entropy reflection anchors and optimizes a chain-masked finite-window KL surrogate for downstream visual dependence. Experiments on reasoning-intensive and general-domain benchmarks show that RAPO delivers substantial gains over strong baselines across multiple LVLM backbones. Mechanism analyses further indicate that reflection anchors are enriched for visually sensitive decision points and that RAPO increases contrastive visual-dependence signals along generated trajectories.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。