用推理可迁移性提升多模态大模型持续学习效果
Reasoning Portability: Guiding Continual Learning for MLLMs in the RLVR Era

- 基于推理层面的可迁移性动态调节正则化强度
- 在分布外样本上保持稳定性能,较基线提升12.0%准确率
- 适合追求持续学习中知识保留与新任务适应的开发者
多模态大语言模型在持续学习中需不断适应新任务并保留旧知识。随着强化学习结合可验证奖励(RLVR)的兴起,需新的指导机制。本文提出推理可迁移性(RP),衡量旧策略行为在新任务上的可复用程度,实验证明推理信号在分布外样本上仍可靠,而答案信号不可靠。据此设计基于推理的动态平衡持续学习(RDB-CL),根据每个样本的RP值调整KL正则化强度:高RP样本使用紧锚点保留可复用推理,低RP样本采用松锚点鼓励探索新路径。实验表明,RDB-CL持续优于基线,在最后准确率上相较原始RLVR提升+12.0%。
原文摘要 · Abstract (English)
Vision-Language Models in Continual Learning (VLM-CL) aim to continuously adapt to new multimodal tasks while retaining prior knowledge. The emerging paradigm that couples Multimodal Large Language Models (MLLMs) with Reinforcement Learning with Verifiable Rewards (RLVR) calls for a new pattern to guide continual adaptation. Advances in reasoning capability now make it feasible to impose constraints at the reasoning level. We formalize portability, a sample-level measure of how reusable the previous policy's behavior is on a new task, and empirically show that reasoning-level signals remain reliable on out-of-distribution samples while answer-level signals do not. We instantiate this as Reasoning Portability (RP) and propose Reasoning-based Dynamic Balance Continual Learning (RDB-CL), which modulates the per-sample Kullback-Leibler regularization in RLVR according to RP: a tight anchor preserves reusable reasoning on high-RP samples, while a relaxed anchor on low-RP samples permits exploration of new reasoning pathways. Experiments show that RDB-CL consistently outperforms baselines, improving Last accuracy by +12.0% over the vanilla RLVR baseline.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。