arXiv:2606.02322cs.LGcs.AI2026-06

用对抗扰动做几何控制,让大模型持续学习不遗忘

Repurposing Adversarial Perturbations for Continual Learning: From Defense to Active Alignment

论文配图:Repurposing Adversarial Perturbations for Continual Learning: From Defense to Active Alignment
图 1 · 摘自论文原文
  • 将对抗扰动转为几何控制信号,稳定持续学习
  • 遗忘率降低,任务间迁移能力更强,鲁棒性提升
  • 可插件集成到多种持续学习方法中,适用性强

在动态环境中,大语言模型需持续适应新任务,但持续学习常面临遗忘、迁移有限及对对抗扰动敏感等问题。为此,我们提出AdvCL,将对抗扰动重用于几何控制,实现稳定持续适应。AdvCL包含三个可插拔模块:Intra-Smooth通过小幅度对抗扰动促进局部平滑;Proto-Clip采用相似性裁剪防止过度对齐当前任务原型;Inter-Align通过方向性对齐前序任务原型,减少表征差距。实验显示,在标准性能与鲁棒性上均有持续提升,遗忘更少,迁移更强。进一步分析表明,Intra-Smooth对扰动设置敏感,Inter-Align显著影响任务相似性与几何距离。三模块组合提供互补增益,且各自可独立嵌入重放、正则化与动态架构等各类持续学习范式,形成通用几何控制机制。

原文摘要 · Abstract (English)

In dynamic environments, large language models need to keep adapting to new tasks, but continual learning often suffers from forgetting, limited transfer, and vulnerability to adversarial perturbations. To address this, we present AdvCL, which repurposes adversarial perturbations as a geometric control signal for stable continual adaptation. AdvCL combines three plug-in modules: Intra-Smooth promotes local smoothness via small adversarial perturbations; Proto-Clip uses similarity clipping to prevent excessive alignment to current task prototype; and Inter-Align applies directional alignment toward previous task prototype to reduce representational gaps. Experiments show consistent gains in both standard performance and robustness, with lower forgetting and stronger transfer. We further analyze key mechanisms by quantifying the sensitivity of Intra-Smooth to perturbation settings and the effect of Inter-Align on task similarity and geometric distance. In summary, the modules provide complementary gains when combined, and each can also be integrated individually into diverse CL paradigms, including replay, regularization, and dynamic architectures, thereby offering a geometric control mechanism for continual learning.

持续学习对抗扰动大模型几何控制

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。