防御持续学习中的后门攻击,通过净化样本和选择性恢复提升模型鲁棒性。
Robust Dynamic Expansion for Continual Learning under Backdoor Attacks via Purification and Selective Recovery

- 利用特征空间语义差异识别可疑样本,实现样本净化。
- 通过伪标签修正与梯度一致性评估,选择性恢复有信息量的中毒样本。
- 构建扰动感知类别原型,保障输入异常时专家路由可靠性。
持续学习(CL)使模型能够从顺序到达的任务中获取新知识,同时保留先前学习的知识。然而,在实际场景中,来自不可信源的任务流可能包含被后门污染的样本,严重威胁持续学习者的稳定性、可塑性和安全性。本文研究了一种称为持续学习下的后门攻击(CLUBA)的挑战性设置,其中每个增量任务可能包含少量恶意操纵的训练样本。不同于传统的持续学习或后门防御场景,CLUBA要求模型在序列更新过程中同时缓解灾难性遗忘、保持适应能力并防止恶意监督的吸收。为此,我们提出一种鲁棒的动态扩展框架,将样本净化、选择性恢复和鲁棒专家路由整合到统一的持续学习范式中。具体地,引入双原型净化(BPP)通过特征空间中的语义差异识别可疑样本。基于净化后的数据,基于梯度差异的鲁棒优化(GDBRO)通过伪标签校正和梯度一致性评估选择性恢复有信息量的中毒样本,既提升了鲁棒性又保持了模型可塑性。此外,基于鲁棒特征一致性的专家选择(RFCBES)构建扰动感知的类别原型,实现在输入被污染或偏移情况下的可靠专家路由。
原文摘要 · Abstract (English)
Continual learning (CL) enables models to acquire new knowledge from sequentially arriving tasks while retaining previously learned knowledge. However, in practical scenarios, task streams collected from untrusted sources may contain backdoor-poisoned samples, posing a critical challenge to the stability, plasticity, and security of continual learners. In this work, we investigate a challenging setting termed Continual Learning Under Backdoor Attack (CLUBA), where each incremental task may involve a small proportion of maliciously manipulated training samples. Unlike conventional continual learning or backdoor defense scenarios, CLUBA requires models to simultaneously mitigate catastrophic forgetting, preserve adaptation capability, and prevent the absorption of malicious supervision during sequential updates. To address this challenge, we propose a robust dynamic-expansion framework that integrates sample purification, selective recovery, and robust expert routing into a unified continual learning paradigm. Specifically, we introduce Bi-Prototype Purification (BPP) to identify suspicious samples by exploiting semantic discrepancies in feature space. Based on purified data, Gradient Discrepancy-based Robustness Optimization (GDBRO) selectively recovers informative poisoned samples through pseudo-label correction and gradient consistency evaluation, improving robustness while maintaining model plasticity. Furthermore, Robust Feature Consistency-based Expert Selection (RFCBES) constructs perturbation-aware class prototypes to enable reliable expert routing under corrupted or shifted inputs.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。