arXiv:2505.15594cs.LGcs.AI2025-05中稿 · the 33rd European …被引 1

测试扩散去噪平滑在多任务下的安全与性能权衡,发现高噪声会严重降效。

Beyond Classification: Evaluating Diffusion Denoised Smoothing for Security-Utility Trade off

  • 用预训练扩散模型预处理输入,提升对抗鲁棒性
  • 高噪声设置使性能下降最高达57%,低噪声防护不足
  • 提出新攻击方法可突破低噪声防御,适合安全研究者参考

尽管基础模型在多种任务中表现优异,但仍易受对抗输入影响。当前研究探索多种增强鲁棒性的方法,其中扩散去噪平滑表现出显著潜力。该方法利用预训练扩散模型在推理前对输入进行预处理。然而,其有效性在分类任务之外仍缺乏系统评估。本文在三个数据集上,针对四种下游任务,采用三种不同对抗攻击算法进行分析。结果表明,基础模型对常规变换具有鲁棒性,但对无失真的干净图像施加高噪声扩散去噪会导致性能最高下降57%。低噪声设置虽能保持性能,却无法有效抵御所有攻击类型。此外,我们提出一种新型攻击策略,专门针对扩散过程本身,可绕过低噪声防御。研究揭示,对抗鲁棒性与性能之间的权衡仍是亟待解决的挑战。

原文摘要 · Abstract (English)

While foundation models demonstrate impressive performance across various tasks, they remain vulnerable to adversarial inputs. Current research explores various approaches to enhance model robustness, with Diffusion Denoised Smoothing emerging as a particularly promising technique. This method employs a pretrained diffusion model to preprocess inputs before model inference. Yet, its effectiveness remains largely unexplored beyond classification. We aim to address this gap by analyzing three datasets with four distinct downstream tasks under three different adversarial attack algorithms. Our findings reveal that while foundation models maintain resilience against conventional transformations, applying high-noise diffusion denoising to clean images without any distortions significantly degrades performance by as high as 57%. Low-noise diffusion settings preserve performance but fail to provide adequate protection across all attack types. Moreover, we introduce a novel attack strategy specifically targeting the diffusion process itself, capable of circumventing defenses in the low-noise regime. Our results suggest that the trade-off between adversarial robustness and performance remains a challenge to be addressed.

扩散模型对抗攻击安全权衡

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。