arXiv:2608.00543cs.CRcs.AI2026-08

黑客可植入隐蔽后门,绕过扩散模型的语义水印检测。

Robust Watermarks Meet Backdoored Models: Evading Diffusion Semantic Watermarks via Stealthy Backdoor

论文配图:Robust Watermarks Meet Backdoored Models: Evading Diffusion Semantic Watermarks via Stealthy Backdoor
图 1 · 摘自论文原文
  • 用频谱正则化构建通用触发器,隐蔽植入编码器后门。
  • 良性图像水印检测率94.4%,触发后逃逸成功率94.6%。
  • 对主流防御方法均有效,适用于安全评估与对抗研究。

尽管语义水印被视为潜在保护生成图像的方案,但其检测流程依赖神经网络,带来了尚未被充分探索的后门攻击风险。本文提出GhostVAE,在变分自编码器(VAE)编码器中植入隐蔽后门,实现对语义水印的有效规避。该方法分两阶段:首先通过功率谱正则化构建通用触发器以增强鲁棒性,再以参数对齐目标训练带后门的VAE编码器。在三种先进语义水印方案和三种主流潜空间扩散模型上进行广泛评估,结果显示,良性图像下水印检测平均真阳性率达94.4%,而触发激活时攻击成功率平均达94.6%。此外,我们系统分析了十七种代表性防御,证明GhostVAE在输入、参数与隐空间均保持隐蔽性。本工作从根本上动摇了语义水印系统的可信度,强调水印部署需考虑端到端安全性,尤其是神经网络组件。

原文摘要 · Abstract (English)

Although semantic watermarking is considered a promising safeguard for images generated by Latent Diffusion Models (LDMs), the reliance of the watermark detection pipeline on neural networks introduces a critical yet underexplored backdoor attack surface. To systematically study this vulnerability, we propose GhostVAE to plant a stealthy backdoor into the encoder of Variational Autoencoder (VAE), enabling reliable evasion of watermark detection. GhostVAE operates in two stages: it first constructs a universal trigger via power spectrum regularization to improve the trigger robustness, and then trains a backdoored VAE encoder with a parameter-aligned objective. Through extensive evaluations across three state-of-the-art semantic watermarking schemes and three widely adopted LDMs, we show that GhostVAE preserves watermark detection performance on benign images (achieving an average true positive rate of 94.4%), while simultaneously enabling highly effective evasion under trigger activation (achieving an average attack success rate of 94.6%). Moreover, we comprehensively analyze seventeen representative defenses and demonstrate that GhostVAE remains stealthy across the input space, parameter space, and latent space. Our work fundamentally undermines the trustworthiness of semantic watermarking systems and highlights that secure deployment of semantic watermarks requires end-to-end security considerations, particularly for neural network components.

语义水印后门攻击扩散模型安全评估

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。