提出防御持续学习中数据投毒攻击的理论框架,揭示了攻击与防御的极限条件。
Theory of Continual Learning Against Data Poisoning Attacks

- 将攻防过程建模为在线零和博弈,分析正则化持续学习中的对抗机制。
- 证明当攻击者污染线性比例任务时,任何防御均无效;但小频率或有限噪声攻击可防御。
- 设计任务验证机制和鲁棒防御算法,实证验证理论结果的有效性。
持续学习(CL)在大语言模型和图像识别等关键领域日益广泛应用,但极易受到数据投毒攻击,导致学习发散或过度风险。尽管威胁显著,现有研究仍缺乏针对此类攻击与防御的系统性理论基础。本文构建了一个理论框架,分析基于正则化的持续学习中的策略性攻防行为。通过将对抗关系建模为在线零和博弈,我们首先确立一个根本性能极限:当攻击者以无界噪声或模式偏移污染线性比例的任务时,任何防御手段均无法成功。随后分析两种可能可防御的情形:攻击频率较低,以及每次攻击的噪声有界。对于前者,提出任务间验证机制以检测投毒并减少累积偏差,促进学习收敛;对于后者,推导出一种鲁棒防御方法,最小化模型对有毒特征的敏感性,可严格加速收敛速率。在真实任务上的大量实验进一步验证了理论结果。
原文摘要 · Abstract (English)
Continual learning (CL), where a model is trained on a sequence of data tasks, is increasingly being adopted across key fields such as large language models and image recognition, yet it remains highly vulnerable to data poisoning that triggers learning divergence or severe excess risk. Despite these threats, a principled theoretical foundation in CL for understanding attack and defense remains lacking. In this paper, we develop a theoretical framework to analyze strategic attacks and defenses in regularization-based CL, a cornerstone of recent CL theory. By framing the adversary-defender interaction as an online zero-sum game, we first establish a fundamental performance limit: no defense succeeds when an adversary poisons a linear proportion of tasks by injecting unbounded noise or pattern shifts in regularization-based CL. We then analyze two possibly defensible scenarios: infrequent attacks and bounded noise per attack. For the former regime, we propose a task-to-task verification mechanism to detect data poisoning and reduce cumulative bias for learning convergence. For the latter regime, we derive a robust defense that minimizes the model's sensitivity to poisoned features, provably accelerating the convergence rate. Extensive experiments on realistic tasks further validate our theoretical results.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。