arXiv:2605.09640cs.CVcs.LG2026-05被引 1

用强化学习优化策略,缓解视觉持续学习中的遗忘问题

Overcoming Catastrophic Forgetting in Visual Continual Learning with Reinforcement Fine-Tuning

论文配图:Overcoming Catastrophic Forgetting in Visual Continual Learning with Reinforcement Fine-Tuning
图 1 · 摘自论文原文
  • 通过轨迹级奖励设计,引导模型保留旧知识
  • 在5种视觉持续学习任务中显著降低遗忘率
  • 适合需要长期记忆的智能系统研发者

近期研究发现,强化微调(RFT)比监督微调(SFT)更抗灾难性遗忘。然而,RFT(如GRPO)在类增量学习(CIL)和域增量学习(DIL)等挑战性视觉持续学习场景中是否有效仍不明确。我们通过初步实验确认,尽管RFT始终优于SFT,但仍存在明显遗忘。我们实证发现根本瓶颈在于轨迹级漂移无关性:在获得相同任务奖励的候选轨迹中,与前一任务策略的KL散度差异显著,且与序列任务中的遗忘强相关。为此,我们提出感知保留的策略优化(RaPO),一种简单有效的RFT方法,通过轨迹级奖励重塑显式缓解遗忘。核心包含两部分:(1) 保留奖励,将轨迹分布漂移转化为连续奖励信号,优先强化各组内知识保留的轨迹;(2) 跨任务优势归一化(CTAN),在任务边界维持奖励统计的指数移动平均,稳定持续学习过程。利用多模态大模型(MLLM)的自由文本泛化能力,我们在五个视觉持续学习设置中全面评估了RaPO。大量实验表明,RaPO表现领先,显著减少灾难性遗忘的同时保持强可塑性。据我们所知,这是首次对视觉持续学习中RFT的系统探索,为未来研究提供启发。

原文摘要 · Abstract (English)

Recent studies suggest that Reinforcement Fine-Tuning (RFT) is inherently more resilient to catastrophic forgetting than Supervised Fine-Tuning (SFT). However, whether RFT (e.g., GRPO) can effectively overcome forgetting in challenging visual continual learning settings, such as class-incremental learning (CIL) and domain-incremental learning (DIL), remains an open problem. Through a pilot study, we confirm that while RFT consistently outperforms SFT, it still suffers from non-negligible forgetting. We empirically trace this bottleneck to Trajectory-level Drift Agnosticism: among candidate rollouts achieving identical task rewards, the KL divergence from the preceding-task policy varies substantially, which strongly correlates with catastrophic forgetting across sequential tasks. Motivated by this insight, we propose Retention-aware Policy Optimization (RaPO), a simple yet effective RFT method that explicitly mitigates forgetting through trajectory-level reward shaping. Specifically, RaPO comprises two core components: (1) Retention Reward that converts trajectory-level distribution drift into a continuous reward signal, preferentially reinforcing knowledge-preserving rollouts within each group; (2) Cross-Task Advantage Normalization (CTAN), which maintains a persistent exponential moving average of reward statistics across task boundaries to stabilize the optimization progress during continual learning. Leveraging the free-form textual generalization of MLLMs, we comprehensively evaluate RaPO across five visual continual learning settings. Extensive experiments demonstrate that RaPO achieves leading performance, substantially reducing catastrophic forgetting while preserving strong plasticity. To the best of our knowledge, this work represents the first systematic exploration of RFT in visual continual learning, offering insights that we hope will inspire future research.

持续学习强化学习视觉模型遗忘缓解

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。