arXiv:2607.02643cs.CV2026-07

用双频谱结构在扩散模型隐空间嵌入水印,提升版权保护鲁棒性。

BiSLW: Bi-Spectral Latent Watermarking for Generative Diffusion Models

论文配图:BiSLW: Bi-Spectral Latent Watermarking for Generative Diffusion Models
图 1 · 摘自论文原文
  • 分高低频带学习嵌入水印,利用隐空间固有频率结构。
  • 水印在再生攻击下仍保持近满分准确率,PSNR提升超3dB。
  • 适合需要版权认证的生成式图像应用,计算开销极低。

基于扩散的生成模型虽革新了视觉内容合成,却面临未经授权使用和缺乏可靠归属的问题。现有水印技术多将隐变量视为静态空间特征图或依赖像素域修改,且未显式利用隐空间的内部频率结构进行双频带冗余嵌入,因而易受扩散过程随机性和再生攻击影响。本文提出可训练的双频谱隐空间水印框架BiSLW,通过学习编码器与解码器,在解码后的扩散隐空间互补频带中联合嵌入对齐的身份信号。低频成分编码全局语义,高频成分捕捉精细纹理,水印分别注入两频带并重组后参与生成,使其成为生成轨迹的内在部分。双频解码器分别从各频带恢复水印,并通过跨频带一致性约束确保语义与纹理嵌入对齐。实验表明,BiSLW在感知保真度与鲁棒性间取得良好平衡,相比先前隐空间水印方法提升超过3 dB的PSNR,且在激进再生和常见失真下保持近乎完美的位准确率,计算开销可忽略不计。

原文摘要 · Abstract (English)

Diffusion-based generative models have transformed visual content synthesis, yet they remain vulnerable to unauthorized usage and lack reliable attribution methods. Existing watermarking techniques often treat latent tensors as static spatial feature maps or depend on pixel-domain modification, and most do not explicitly leverage the internal frequency structure of the latent space for dual-band redundant embedding, leaving them susceptible to the stochastic nature of diffusion and regeneration attacks. We introduce BiSLW, a trainable bi-spectral latent watermarking framework that jointly embeds aligned identity signals across complementary spectral bands of the decoded diffusion latent using learned encoders and decoders, going beyond fixed-pattern frequency approaches. We leverage the inherent frequency structure of diffusion latents to design a dual-band watermarking framework. Low-frequency components encode global semantics, while high-frequency components capture fine texture. We exploit this structure to embed watermarks across complementary spectral bands. The watermark is independently injected into both bands via learned encoders and recombined before decoding, ensuring it becomes intrinsic to the generative trajectory. Dual spectral decoders recover the watermark from each band, while a cross-band consistency constraint enforces alignment between semantic and textural embeddings. Experiments show that BiSLW achieves a strong balance between perceptual fidelity and robustness, improving PSNR by over 3 dB compared to prior latent diffusion watermarking methods while preserving near-perfect bit accuracy under aggressive regeneration and common distortions, all with negligible computational overhead.

水印扩散模型隐空间版权保护

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。