揭示多领域强化学习中干扰的局部机制并实现精准恢复。
A Local Perturbation Theory for Cross-Domain Interference and Recovery in Multi-Domain RL

- 基于局部扰动理论,发现干扰源于共享计算路径上的方向冲突。
- 数学领域微调后性能从57.66提升至66.04,平均分达66.39。
- 无需训练即可通过稀疏代理坐标回滚,验证损伤定位性。
强化学习(RL)微调能提升大语言模型在数学推理、代码生成、问答和创意写作等单一领域的表现,但某一领域的训练常导致其他领域性能下降。现有灾难性遗忘或全局梯度冲突解释不完整:即使全模型梯度几乎正交,仍存在显著干扰。我们发现单领域RL产生稀疏且小幅度的参数更新,不同领域虽仅部分重叠关键神经元,却共享大量活跃计算路径,其更新方向决定协同或冲突。基于此,我们在局部扰动模型下证明:后期领域训练对早期领域的主要损害来自二阶损伤项,该效应在低维共享冲突子空间集中。简短的领域刷新可压缩此有害成分,实现选择性恢复且副作用小。实验显示,在代码→数学→问答→创意写作流程后,仅一次数学刷新即可将数学得分从57.66提升至66.04,平均分达66.39。此外,对数学-问答对的稀疏代理冲突坐标集进行无训练回滚,部分恢复数学性能,为局部损伤提供直接证据。结果为多领域强化学习中的干扰与恢复提供了局部化机理解释。
原文摘要 · Abstract (English)
Reinforcement learning (RL) post-training improves large language models (LLMs) on individual domains such as mathematical reasoning, code generation, question answering, and creative writing (CW), but training on one domain often degrades performance on others. Existing explanations based on catastrophic forgetting or global gradient conflict are incomplete: substantial interference can occur even when full-model gradients are nearly orthogonal. We show that single-domain RL produces sparse, small-magnitude parameter edits with weak overlap among top-changed neurons, while different domains still share substantial active computation routes on which update directions determine whether they act synergistically or conflict. Guided by this observation, we prove under a local perturbation model of multi-domain RL that later-domain training harms an earlier domain mainly through a second-order damage term, which under the observed sparse route structure concentrates in a low-dimensional shared conflict subspace. Moreover, a short domain refresh contracts the harmful component on this subspace, enabling selective recovery with limited collateral damage. Consistent with the theory, a brief Re-Math refresh after Code $\rightarrow$ Math $\rightarrow$ QA $\rightarrow$ CW recovers Math from 57.66 to 66.04 while largely preserving performance on the other domains, yielding the best average score of 66.39. Beyond refresh, a training-free rollback on a sparse proxy conflict coordinate set for the Math-QA pair partially restores Math, providing direct proxy-level evidence for localized damage. These results provide a localized mechanistic account of interference and recovery in multi-domain RL.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。