arXiv:2510.01968cs.SDcs.LG2025-10被引 1

无需训练即可实现高鲁棒性音频水印,隐匿性强且效果稳定。

Multi-bit Audio Watermarking

  • 通过优化预训练音频VAE的潜在空间,添加不可察觉的扰动来嵌入水印。
  • 在MUSDB18-HQ数据集上,16位水印平均误码率优于AudioSeal等现有方法。
  • 适用于音乐版权保护,特别适合对隐蔽性和鲁棒性要求高的场景。

我们提出Timbru,一种后处理式音频水印模型,在无需训练嵌入-检测器模型的前提下,实现了当前最优的鲁棒性与不可察觉性平衡。给定任意44.1 kHz立体声音乐片段,该方法通过对预训练音频变分自编码器(VAE)的潜在空间进行逐音频梯度优化,基于消息损失与感知损失的联合引导,添加不可察觉的扰动以嵌入水印。水印可通过预训练的CLAP模型提取。我们在MUSDB18-HQ数据集上评估了16位水印在常见攻击(滤波、噪声、压缩、重采样、裁剪、再生)下的表现,结果表明其平均比特误码率最低,同时保持良好的听觉质量,展示了一条高效、无需数据集依赖的不可察觉音频水印路径。

原文摘要 · Abstract (English)

We present Timbru, a post-hoc audio watermarking model that achieves state-of-the-art robustness and imperceptibility trade-offs without training an embedder-detector model. Given any 44.1 kHz stereo music snippet, our method performs per-audio gradient optimization to add imperceptible perturbations in the latent space of a pretrained audio VAE, guided by a combined message and perceptual loss. The watermark can then be extracted using a pretrained CLAP model. We evaluate 16-bit watermarking on MUSDB18-HQ against AudioSeal, WavMark, and SilentCipher across common filtering, noise, compression, resampling, cropping, and regeneration attacks. Our approach attains the best average bit error rates, while preserving perceptual quality, demonstrating an efficient, dataset-free path to imperceptible audio watermarking.

音频水印隐写VAECLAP

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。