发现近期自适应攻击存在实现错误,修复后防御仍有效且攻击更贴近人类感知。
A Note on Implementation Errors in Recent Adaptive Attacks Against Multi-Resolution Self-Ensembles
- 修正攻击中超出原定 $L_ty = 8/255$ 边界的实现错误
- 修复后防御在 $L_ty = 8/255$ 下仍保持显著鲁棒性
- 正确约束的攻击与人类视觉感知更一致,提示需重新思考评估标准
本文指出近期针对多分辨率自集成防御(Fort and Lakshminarayanan [2024])的自适应攻击(Zhang et al. [2024])存在实现问题:攻击扰动超出标准 $L_\infty = 8/255$ 约20倍,最高达 $L_\infty = 160/255$。当攻击被正确限制在原定范围内时,该防御仍表现出非平凡的鲁棒性。分析还揭示,经过合理约束的自适应攻击往往与人类感知高度一致,这提示我们应重新审视当前对抗鲁棒性的衡量方式。
原文摘要 · Abstract (English)
This note documents an implementation issue in recent adaptive attacks (Zhang et al. [2024]) against the multi-resolution self-ensemble defense (Fort and Lakshminarayanan [2024]). The implementation allowed adversarial perturbations to exceed the standard $L_\infty = 8/255$ bound by up to a factor of 20$\times$, reaching magnitudes of up to $L_\infty = 160/255$. When attacks are properly constrained within the intended bounds, the defense maintains non-trivial robustness. Beyond highlighting the importance of careful validation in adversarial machine learning research, our analysis reveals an intriguing finding: properly bounded adaptive attacks against strong multi-resolution self-ensembles often align with human perception, suggesting the need to reconsider how we measure adversarial robustness.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。