利用公开VAE破解扩散模型水印,效果远超已有攻击方法。
A Crack in the Bark: Leveraging Public Knowledge to Remove Tree-Ring Watermarks
- 仅需公开的变分自编码器,即可逼近模型中间特征空间。
- 水印检测AUC从0.993降至0.153,图像质量仍保持良好。
- 揭示行业忽视的公共组件风险,适合安全与模型保护研究者。
本文提出一种针对扩散模型水印技术Tree-Ring的新攻击方法,该技术以高隐蔽性和强抗移除性著称。不同于以往依赖强大攻击能力的假设,本攻击只需访问目标扩散模型训练所用的变分自编码器(VAE),而此类组件常被公开。通过利用公开的VAE,攻击者可有效逼近模型中间潜在空间,进而实施更高效的代理攻击。实验表明,该方法使Tree-Ring检测器的ROC曲线AUC从0.993降至0.153,PR曲线AUC从0.994降至0.385,同时保持高图像质量。值得注意的是,该攻击性能优于假设完全访问扩散模型的现有方法。结果揭示了当前工业实践中重复使用公开VAE训练扩散模型带来的安全隐患,且指出此前被忽视的检测精度指标在真实部署中不达标。
原文摘要 · Abstract (English)
We present a novel attack specifically designed against Tree-Ring, a watermarking technique for diffusion models known for its high imperceptibility and robustness against removal attacks. Unlike previous removal attacks, which rely on strong assumptions about attacker capabilities, our attack only requires access to the variational autoencoder that was used to train the target diffusion model, a component that is often publicly available. By leveraging this variational autoencoder, the attacker can approximate the model's intermediate latent space, enabling more effective surrogate-based attacks. Our evaluation shows that this approach leads to a dramatic reduction in the AUC of Tree-Ring detector's ROC and PR curves, decreasing from 0.993 to 0.153 and from 0.994 to 0.385, respectively, while maintaining high image quality. Notably, our attacks outperform existing methods that assume full access to the diffusion model. These findings highlight the risk of reusing public autoencoders to train diffusion models -- a threat not considered by current industry practices. Furthermore, the results suggest that the Tree-Ring detector's precision, a metric that has been overlooked by previous evaluations, falls short of the requirements for real-world deployment.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。