音乐生成中首次实现内容级水印,抗重编码攻击。
MusicMark: A Robust Generative Watermarking Framework for Music Generation

- 在扩散模型的语义潜空间嵌入水印,与生成过程一体化。
- 对神经编解码重合成攻击等多类攻击保持高鲁棒性。
- 适合需版权溯源的AI音乐创作与平台应用。
AI音乐生成技术快速发展,催生了可靠的溯源水印需求。然而现有音频水印研究主要针对语音,难以直接应用于结构复杂、声学纹理丰富的音乐。多数方法为后处理式,在生成后添加不可察觉扰动,易受变换及神经编解码重合成攻击影响,且因生成与水印分离,可能被跳过。为此,我们提出MusicMark,据知是首个面向音乐生成的生成式水印框架。该框架在扩散模型的去噪过程中,通过引入水印适配器,将水印信息嵌入语义潜空间,使其成为音乐内容的一部分。适配器与检测器通过联合目标训练,既约束水印潜变量接近未水印参考值以保证音质,又通过攻击增强提升鲁棒性。实验表明,MusicMark在多种攻击(包括神经编解码重合成)下显著优于后处理基线,同时保持相当的生成质量。我们还提出一种伴奏覆盖攻击,保留音乐内容而转换人声,结果显示MusicMark仍比后处理方法更鲁棒。
原文摘要 · Abstract (English)
AI music generation has rapidly advanced alongside commercial platforms, raising the need for reliable watermarking for provenance and attribution. However, existing audio watermarking research has largely focused on speech, and applying speech-oriented methods to music is challenging due to music's complex structure and rich acoustic texture. Most existing methods are post-hoc, adding imperceptible perturbations after generation rather than embedding watermarks as part of the content. This makes them fragile under transformations and especially vulnerable to neural codec re-synthesis, which can discard imperceptible residual signals. Moreover, since generation and watermarking are decoupled, the watermarking step can be bypassed or omitted, weakening provenance guarantees. To address these issues, we propose MusicMark, which, to the best of our knowledge, is the first generative watermarking framework for music. Specifically, MusicMark embeds watermark messages into the semantic latent space during generation, incorporating the watermark as part of the musical content and ensuring robustness against diverse attacks, particularly neural codec re-synthesis. To this end, we introduce a watermark adapter into a diffusion-based generation model to embed watermark messages across denoising steps. The adapter and detector are trained with a joint objective that preserves fidelity by constraining watermarked latents close to their unwatermarked reference latents, while improving robustness through attack augmentations. Experiments demonstrate that MusicMark substantially outperforms post-hoc baselines across diverse attacks including neural codec re-synthesis, while maintaining comparable generation quality. We further introduce a cover-song attack, converting the singing voice while preserving musical content, and show that MusicMark remains more robust than post-hoc methods.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。