arXiv:2512.10402cs.LGcs.AI2025-12KDD

提出可解释的鲁棒后门攻击框架,用理论指导低毒率高效攻击。

The Eminence in Shadow: Exploiting Feature Boundary Ambiguity for Robust Backdoor Attacks

  • 基于决策边界模糊性设计隐蔽触发器,实现极低污染率攻击。
  • 攻击成功率超90%,毒样本率低于0.1%,清洁准确率损失微乎其微。
  • 理论可证,适用于多种模型与数据集,适合研究安全漏洞者。

深度神经网络在关键应用中广泛应用,却易受后门攻击威胁,传统方法依赖经验性暴力搜索。尽管后门研究已取得显著实证进展,但缺乏严谨的理论分析,限制了对攻击机制的理解,制约了攻击的可预测性与适应性。为此,本文提供针对后门攻击的理论分析,聚焦稀疏决策边界如何导致模型被显著操控。基于此发现,推导出闭式表达的模糊边界区域,在该区域内,极少数量的标签篡改样本即可引发严重误分类。影响函数分析进一步量化了这些边缘样本引起的显著参数偏移,同时对干净准确率影响极小,从理论上解释了为何极低毒率即可实现高效攻击。据此,我们提出Eminence——一种可解释且鲁棒的黑盒后门框架,具备理论保证和内在隐蔽性。Eminence优化出通用、视觉上隐蔽的触发器,战略性地利用脆弱决策边界,在极低毒率(<0.1%)下实现稳健误分类,远优于当前最优方法(通常需>1%)。全面实验验证了理论讨论,并证实毒力与边界操纵间存在指数关系。Eminence保持>90%攻击成功率,清洁准确率损失可忽略,且在多种模型、数据集与场景间具有高迁移性。

原文摘要 · Abstract (English)

Deep neural networks (DNNs) underpin critical applications yet remain vulnerable to backdoor attacks, typically reliant on heuristic brute-force methods. Despite significant empirical advancements in backdoor research, the lack of rigorous theoretical analysis limits understanding of underlying mechanisms, constraining attack predictability and adaptability. Therefore, we provide a theoretical analysis targeting backdoor attacks, focusing on how sparse decision boundaries enable disproportionate model manipulation. Based on this finding, we derive a closed-form, ambiguous boundary region, wherein negligible relabeled samples induce substantial misclassification. Influence function analysis further quantifies significant parameter shifts caused by these margin samples, with minimal impact on clean accuracy, formally grounding why such low poison rates suffice for efficacious attacks. Leveraging these insights, we propose Eminence, an explainable and robust black-box backdoor framework with provable theoretical guarantees and inherent stealth properties. Eminence optimizes a universal, visually subtle trigger that strategically exploits vulnerable decision boundaries and effectively achieves robust misclassification with exceptionally low poison rates (< 0.1%, compared to SOTA methods typically requiring > 1%). Comprehensive experiments validate our theoretical discussions and demonstrate the effectiveness of Eminence, confirming an exponential relationship between margin poisoning and adversarial boundary manipulation. Eminence maintains > 90% attack success rate, exhibits negligible clean-accuracy loss, and demonstrates high transferability across diverse models, datasets and scenarios.

后门攻击理论分析隐蔽触发器低毒率

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。