arXiv:2503.06966cs.CV2025-03

提出首个针对去噪模型的语义扰动攻击,让图像看起来干净却含义被篡改。

MIGA: Mutual Information-Guided Attack on Denoising Models for Semantic Manipulation

  • 通过最小化原始图与去噪图的互信息,破坏语义保留能力。
  • 在5个数据集上测试,4种去噪模型均出现显著语义失真。
  • 揭示去噪模型存在隐蔽安全漏洞,适合关注模型鲁棒性的研究者。

基于深度学习的去噪模型广泛应用于视觉任务中,作为滤波器去除噪声的同时保留关键语义信息,并在防御对抗性扰动方面发挥重要作用。然而,由于依赖特定噪声假设,这些模型可能内在地易受对抗攻击。现有攻击主要损害视觉清晰度,忽略语义操控,导致易检测或效果有限。本文提出互信息引导攻击(MIGA),是首个直接针对去噪模型的语义操纵攻击方法,通过对抗扰动策略性破坏其保持语义内容的能力。通过最小化原始图像与去噪图像之间的互信息(衡量语义相似性),MIGA迫使去噪器生成看似清晰但语义已扭曲的输出。这些图像虽视觉上合理,却编码了系统性语义偏差,且在去噪输出中持续存在,可通过下游任务性能定量评估。我们提出新的评估指标,在四种去噪模型和五个数据集上系统评估MIGA,证明其在破坏语义保真度方面的持续有效性。研究结果表明,去噪模型并非始终稳健,可能在实际应用中引入安全风险。

原文摘要 · Abstract (English)

Deep learning-based denoising models have been widely employed in vision tasks, functioning as filters to eliminate noise while retaining crucial semantic information. Additionally, they play a vital role in defending against adversarial perturbations that threaten downstream tasks. However, these models can be intrinsically susceptible to adversarial attacks due to their dependence on specific noise assumptions. Existing attacks on denoising models mainly aim at deteriorating visual clarity while neglecting semantic manipulation, rendering them either easily detectable or limited in effectiveness. In this paper, we propose Mutual Information-Guided Attack (MIGA), the first method designed to directly attack deep denoising models by strategically disrupting their ability to preserve semantic content via adversarial perturbations. By minimizing the mutual information between the original and denoised images, a measure of semantic similarity. MIGA forces the denoiser to produce perceptually clean yet semantically altered outputs. While these images appear visually plausible, they encode systematically distorted semantics, revealing a fundamental vulnerability in denoising models. These distortions persist in denoised outputs and can be quantitatively assessed through downstream task performance. We propose new evaluation metrics and systematically assess MIGA on four denoising models across five datasets, demonstrating its consistent effectiveness in disrupting semantic fidelity. Our findings suggest that denoising models are not always robust and can introduce security risks in real-world applications.

对抗攻击去噪模型语义安全互信息

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。