arXiv:2506.13024cs.CRcs.LG2025-06ICML被引 4

认证鲁棒性不等于模型安全,警惕虚假保障陷阱

Position: Certified Robustness Does Not (Yet) Imply Model Security

  • 指出认证鲁棒性存在检测但无法区分攻击的悖论
  • 揭露用户对'保证'鲁棒性的误解与实际能力脱节
  • 呼吁建立可操作标准,推动认证技术落地应用

尽管认证鲁棒性被广泛视为应对人工智能系统中对抗样本的解决方案,但在实际应用前仍面临诸多挑战。本文揭示当前研究中的关键缺口:检测却无法区分攻击的悖论、从业者评估认证方案缺乏明确标准,以及用户对‘保证’鲁棒性承诺所产生的潜在安全风险。这些因素导致认证的呈现方式与其真实能力之间存在认知错位。本文是一篇立场声明,呼吁认证研究社区采取具体行动,解决这些根本性问题,推动该领域向实际可用迈进。

原文摘要 · Abstract (English)

While certified robustness is widely promoted as a solution to adversarial examples in Artificial Intelligence systems, significant challenges remain before these techniques can be meaningfully deployed in real-world applications. We identify critical gaps in current research, including the paradox of detection without distinction, the lack of clear criteria for practitioners to evaluate certification schemes, and the potential security risks arising from users' expectations surrounding ``guaranteed" robustness claims. These create an alignment issue between how certifications are presented and perceived, relative to their actual capabilities. This position paper is a call to arms for the certification research community, proposing concrete steps to address these fundamental challenges and advance the field toward practical applicability.

鲁棒性安全评估对抗样本

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。