提升深度模型防御自适应后门攻击的能力,识别隐蔽恶意触发器。
TED-LaST: Towards Robust Backdoor Defense Against Adaptive Attacks
- 通过标签监督动态追踪与自适应层加权增强拓扑检测鲁棒性。
- 在CIFAR-10、GTSRB和ImageNet100上对多种攻击实现95%以上检测率。
- 适合关注模型安全与对抗训练的工程师及研究人员。
深度神经网络易受后门攻击,攻击者在训练中植入隐藏触发器以操控模型行为。拓扑演化动力学(TED)是近年检测此类攻击的有效工具,但对自适应攻击导致的层间拓扑分布扭曲仍显脆弱。为此,我们提出TED-LaST(针对洗衣、慢释放和目标映射攻击策略的拓扑演化动力学),引入标签监督动态追踪与自适应层强调两项创新,可有效识别传统TED难以发现的隐蔽威胁,即便在拓扑空间不可分或存在微弱扰动时亦然。我们系统梳理并分类了现有自适应攻击的数据投毒手法,提出具备目标映射能力的增强型自适应攻击,可动态转移恶意任务并最大化隐蔽性。在多个数据集(CIFAR-10、GTSRB、ImageNet100)和模型架构(ResNet20、ResNet101)上的全面实验表明,TED-LaST能有效抵御Adap-Blend、Adapt-Patch及所提增强攻击,显著提升深层模型对演进式威胁的安全防护能力。
原文摘要 · Abstract (English)
Deep Neural Networks (DNNs) are vulnerable to backdoor attacks, where attackers implant hidden triggers during training to maliciously control model behavior. Topological Evolution Dynamics (TED) has recently emerged as a powerful tool for detecting backdoor attacks in DNNs. However, TED can be vulnerable to backdoor attacks that adaptively distort topological representation distributions across network layers. To address this limitation, we propose TED-LaST (Topological Evolution Dynamics against Laundry, Slow release, and Target mapping attack strategies), a novel defense strategy that enhances TED's robustness against adaptive attacks. TED-LaST introduces two key innovations: label-supervised dynamics tracking and adaptive layer emphasis. These enhancements enable the identification of stealthy threats that evade traditional TED-based defenses, even in cases of inseparability in topological space and subtle topological perturbations. We review and classify data poisoning tricks in state-of-the-art adaptive attacks and propose enhanced adaptive attack with target mapping, which can dynamically shift malicious tasks and fully leverage the stealthiness that adaptive attacks possess. Our comprehensive experiments on multiple datasets (CIFAR-10, GTSRB, and ImageNet100) and model architectures (ResNet20, ResNet101) show that TED-LaST effectively counteracts sophisticated backdoors like Adap-Blend, Adapt-Patch, and the proposed enhanced adaptive attack. TED-LaST sets a new benchmark for robust backdoor detection, substantially enhancing DNN security against evolving threats.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。