动态调整正则化,让隐秘文生图后门攻击更有效且保真。
Beyond the False Trade-off: Adaptive EWC for Stealthy and Generalizable T2I Backdoors

- 用动态调节的EWC正则化,替代固定参数的保守训练。
- 在弱触发器下仍保持高攻击成功率和图像保真度。
- 适合研究隐秘攻击与模型鲁棒性的安全方向人员。
保持模型保真度对隐蔽的文本到图像(T2I)后门攻击至关重要。现有方法如Learning without Forgetting(LwF)依赖输出级知识蒸馏,正则化能力有限。本文引入基于参数的弹性权重巩固(EWC)作为替代方案。然而,标准静态EWC采用固定正则化权重lambda和均方误差损失,导致攻击成功率(ASR)与保真度之间存在人为权衡,尤其在弱触发器下性能显著下降。为此,我们提出基于余弦感知的自适应EWC,通过余弦相似度计算语义效用并动态调度正则化强度。该方法将EWC从固定惩罚转变为上下文敏感约束,在维持高ASR的同时有效保护模型原始性能。实验表明,本方法在ASR-保真度平衡及跨域(OOD)数据集鲁棒性方面优于现有基线。
原文摘要 · Abstract (English)
Preserving model fidelity is essential for stealthy text-to-image (T2I) backdoor attacks. Existing methods such as Learning without Forgetting (LwF) rely on output-based distillation, which provides limited regularization. We introduce Elastic Weight Consolidation (EWC) as a parameter-based alternative for preserving fidelity in backdoor learning. While stronger in principle, we show that standard static EWC with a fixed regularization weight lambda and mean-squared utility loss creates an artificial trade-off between attack success rate (ASR) and fidelity, particularly degrading performance on weak triggers. To address this, we propose Cosine-Aware Adaptive EWC, which dynamically adjusts EWC regularization using a cosine-based semantic utility and adaptive scheduling. This approach transforms EWC from a fixed penalty into a context-sensitive constraint, maintaining high ASR while preserving model fidelity. Experiments demonstrate improved ASR-fidelity balance and enhanced robustness on out-of-domain (OOD) datasets compared to existing baselines.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。