arXiv:2603.00711cs.CRcs.CV2026-03

用图神经网络生成隐形后门,0.16%污染率下攻破多个类别

IU: Imperceptible Universal Backdoor Attack

  • 用图卷积网络建模类别关系,生成视觉不可见的干扰
  • 在ImageNet上以0.16%污染率实现最高91.3%攻击成功率
  • 可规避主流防御,适合研究模型安全与对抗攻击者

后门攻击对深度神经网络的安全构成重大威胁,但现有通用后门多依赖视觉显著模式,易被检测且难以规模化。本文提出一种新型不可察觉的通用后门攻击方法,能以极低污染率同时控制所有目标类别,并保持隐蔽性。核心思想是利用图卷积网络(GCNs)建模类别间关系,生成既有效又视觉不可见的类特定扰动。所提框架优化双目标损失函数,平衡隐蔽性(以PSNR等感知相似度指标衡量)与攻击成功率(ASR),支持可扩展的多目标后门注入。在ImageNet-1K和ResNet架构上的大量实验表明,本方法在仅0.16%污染率下即可实现高达91.3%的攻击成功率,同时维持良性准确率并避开当前最先进的防御机制。结果凸显了隐形通用后门的新兴风险,亟需更鲁棒的检测与缓解策略。

原文摘要 · Abstract (English)

Backdoor attacks pose a critical threat to the security of deep neural networks, yet existing efforts on universal backdoors often rely on visually salient patterns, making them easier to detect and less practical at scale. In this work, we introduce a novel imperceptible universal backdoor attack that simultaneously controls all target classes with minimal poisoning while preserving stealth. Our key idea is to leverage graph convolutional networks (GCNs) to model inter-class relationships and generate class-specific perturbations that are both effective and visually invisible. The proposed framework optimizes a dual-objective loss that balances stealthiness (measured by perceptual similarity metrics such as PSNR) and attack success rate (ASR), enabling scalable, multi-target backdoor injection. Extensive experiments on ImageNet-1K with ResNet architectures demonstrate that our method achieves high ASR (up to 91.3%) under poisoning rates as low as 0.16%, while maintaining benign accuracy and evading state-of-the-art defenses. These results highlight the emerging risks of invisible universal backdoors and call for more robust detection and mitigation strategies.

后门攻击图像安全图神经网络隐蔽性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。