arXiv:2511.13143cs.CRcs.AI2025-11被引 2

系统评估后门防御方法,揭示评测标准不统一问题

SoK: The Last Line of Defense: On Backdoor Defense Evaluation

  • 梳理183篇论文,分析防御方法的评测方式差异
  • 超3000次实验显示不同设置下防御效果差异显著
  • 建议标准化评测流程,提升防御方案可信度

后门攻击通过植入隐蔽漏洞威胁深度学习模型,虽已有诸多防御方法提出,但评价方法异质性严重,难以公平比较。本文通过对2018至2025年间发表于主要人工智能与安全会议的183篇后门防御论文进行系统性文献综述与实证评估,分析其特性与评测方法。研究发现,实验设置、评估指标与威胁模型假设存在显著不一致。通过在MNIST、CIFAR-100、ImageNet-1K三个数据集上,使用ResNet-18、VGG-19、ViT-B/16、DenseNet-121四种模型架构,针对16种代表性防御和五种常见攻击,完成超过3,000次实验,验证了防御效果在不同评测环境下波动剧烈。识别出当前评测实践中的关键缺陷:计算开销报告不足、良性条件下行为未充分评估、超参数选择存在偏差、实验覆盖不全。基于结果,提出具体挑战与可操作建议,推动未来防御评测的标准化与改进。本工作旨在为研究人员与产业界提供开发、评估与部署防御方案的实用洞见。

原文摘要 · Abstract (English)

Backdoor attacks pose a significant threat to deep learning models by implanting hidden vulnerabilities that can be activated by malicious inputs. While numerous defenses have been proposed to mitigate these attacks, the heterogeneous landscape of evaluation methodologies hinders fair comparison between defenses. This work presents a systematic (meta-)analysis of backdoor defenses through a comprehensive literature review and empirical evaluation. We analyzed 183 backdoor defense papers published between 2018 and 2025 across major AI and security venues, examining the properties and evaluation methodologies of these defenses. Our analysis reveals significant inconsistencies in experimental setups, evaluation metrics, and threat model assumptions in the literature. Through extensive experiments involving three datasets (MNIST, CIFAR-100, ImageNet-1K), four model architectures (ResNet-18, VGG-19, ViT-B/16, DenseNet-121), 16 representative defenses, and five commonly used attacks, totaling over 3\,000 experiments, we demonstrate that defense effectiveness varies substantially across different evaluation setups. We identify critical gaps in current evaluation practices, including insufficient reporting of computational overhead and behavior under benign conditions, bias in hyperparameter selection, and incomplete experimentation. Based on our findings, we provide concrete challenges and well-motivated recommendations to standardize and improve future defense evaluations. Our work aims to equip researchers and industry practitioners with actionable insights for developing, assessing, and deploying defenses to different systems.

后门攻击防御评测安全评估

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。