用强化学习让大模型在长推理中不丢掉视觉信息
Remember-R1: Mitigating Long-Context Visual Forgetting through Reinforcement Learning

- 通过强化学习直接监督推理过程中的视觉使用
- 显著提升多模态模型在长链推理中的准确率
- 适合研究视觉-语言长时依赖的学者参考
多模态大语言模型在处理复杂任务时越来越依赖长链式推理。然而,随着推理序列变长,模型可能逐渐减少对视觉证据的依赖,转而过度依赖累积的文本上下文,导致视觉遗忘问题。现有方法未直接约束视觉信息在原始推理轨迹中的使用与保持,难以有效解决该问题。为此,我们提出 Remember-R1,一种基于强化学习的框架,通过在原始推理轨迹上施加过程级监督,缓解长上下文视觉遗忘。具体而言,Remember-R1 引入奖励机制,鼓励更广泛覆盖匹配的视觉关键词、增强后期推理步骤中对视觉信息的持续依赖,并聚焦于与问题相关的关键图像区域。在多个模型规模和多样化的多模态基准测试中,Remember-R1 均表现出一致的性能提升。进一步分析表明,该方法能有效减缓生成过程中视觉注意力的衰减,验证了其在缓解长上下文视觉遗忘方面的有效性。
原文摘要 · Abstract (English)
Multimodal large language models (MLLMs) increasingly rely on long chain-of-thought reasoning for complex tasks. However, as reasoning sequences lengthen, models may gradually rely less on visual evidence and more on accumulated textual context, leading to visual forgetting. Existing approaches do not directly constrain how visual evidence is used and maintained along the original reasoning trajectory, leaving long-context visual forgetting insufficiently addressed. To address this issue, we propose Remember-R1, a reinforcement learning framework that mitigates long-context visual forgetting by applying process-level supervision directly on the original reasoning trajectory. Specifically, Remember-R1 introduces rewards that encourage broader coverage of matched visual keywords, stronger persistence of visual dependence in later reasoning steps, and greater focus on question-relevant image regions. Experiments across multiple model scales and diverse multimodal benchmarks demonstrate that Remember-R1 consistently improves reasoning performance. Additional analyses further show that it slows the decline of visual attention during generation, supporting its effectiveness in mitigating long-context visual forgetting.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。