提出高效内存的完整梯度攻击框架,精准评估扩散与朗之万净化防御的鲁棒性。
Memory Efficient Full-gradient Attacks (MEFA) Framework for Adversarial Defense Evaluations

- 采用梯度检查点技术,用重计算换内存,实现长净化路径的精确梯度计算。
- 在扩散模型和基于能量的模型上,发现近似梯度评估会高估防御鲁棒性。
- 支持可控随机性评估,适合研究者验证迭代随机防御的真实抗攻击能力。
本文研究白盒对抗攻击下迭代随机净化防御的稳健评估。核心洞察是:通过梯度检查点技术,以额外重计算为代价大幅降低内存消耗,使长净化轨迹的端到端精确梯度计算成为可能。这使得针对基于扩散和朗之万采样的净化防御的完整梯度自适应攻击成为现实,而此前因内存限制常依赖近似反向传播,导致攻击信号削弱并可能高估鲁棒性。同时,迭代净化中的随机性常未受控,不同轨迹会显著影响报告的鲁棒性指标。基于此,我们提出一种内存高效的全梯度评估框架,结合检查点反向传播与可控随机性的评估协议,在缓解内存瓶颈的同时保留精确梯度。实验评估了基于扩散的净化与基于能量模型(EBMs)的朗之万采样,结果表明全梯度攻击能揭示近似评估遗漏的漏洞。该框架实现了当前最优的ℓ∞和ℓ2白盒攻击,并支持对分布外鲁棒性的探测。总体表明,精确梯度评估对于可靠基准测试迭代随机防御至关重要。
原文摘要 · Abstract (English)
This work studies the robust evaluation of iterative stochastic purification defenses under white-box adversarial attacks. Our key technical insight is that gradient checkpointing makes exact end-to-end gradient computation through long purification trajectories practical by trading additional recomputation for substantially lower memory usage. This enables full-gradient adaptive attacks against diffusion- and Langevin-based purification defenses, where prior evaluations often resort to approximate backpropagation due to memory constraints. These approximations can weaken the attack signal and risk overestimating robustness. In parallel, stochasticity in iterative purification is frequently under-controlled, even though different purification trajectories can substantially change reported robustness metrics. Building on this insight, we introduce a memory-efficient full-gradient evaluation framework for stochastic purification defenses. The framework combines checkpointed backpropagation with evaluation protocols that control stochastic variability, thereby reducing memory bottlenecks while preserving exact gradients. We evaluate diffusion-based purification and Langevin sampling with Energy-Based Models (EBMs), demonstrating that full-gradient attacks uncover vulnerabilities missed by approximate-gradient evaluations. Our framework yields stronger state-of-the-art $\ell_{\infty}$ and $\ell_{2}$ white-box attacks and further supports probing out-of-distribution robustness. Overall, our results show that exact-gradient evaluation is essential for reliable benchmarking of iterative stochastic defenses.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。