arXiv:2512.18791cs.SDcs.AI2025-12

为语音生成扩散模型设计通用水印,提升音频质量与抗攻击能力。

Smark: A Watermark for Text-to-Speech Diffusion Models via Discrete Wavelet Transform

  • 基于离散小波变换,在低频区域嵌入水印,适配所有扩散模型。
  • 水印提取准确率超95%,音频失真低于0.1分贝,保持高保真。
  • 适合需要版权保护的语音合成系统,尤其抗逆向扩散攻击。

文本转语音(TTS)扩散模型能生成高质量语音,但对模型知识产权保护和合法使用下的语音溯源带来挑战。音频水印是潜在解决方案,但现有方法因不同模型结构差异,常针对特定模型设计,且损害音质,实用性受限。为此,本文提出一种通用水印方案Smark,通过在所有TTS扩散模型共有的反向扩散范式下,设计轻量级嵌入框架实现水印。为降低对音质影响,Smark利用离散小波变换(DWT)将水印嵌入音频中相对稳定的低频区域,确保水印与音频无缝融合,并抵抗反向扩散过程中的移除攻击。大量实验评估了多种模拟真实攻击场景下的音质与水印性能。结果表明,Smark在音质和水印提取准确性方面均表现优异,水印提取准确率超过95%,音质失真低于0.1 dB。

原文摘要 · Abstract (English)

Text-to-Speech (TTS) diffusion models generate high-quality speech, which raises challenges for the model intellectual property protection and speech tracing for legal use. Audio watermarking is a promising solution. However, due to the structural differences among various TTS diffusion models, existing watermarking methods are often designed for a specific model and degrade audio quality, which limits their practical applicability. To address this dilemma, this paper proposes a universal watermarking scheme for TTS diffusion models, termed Smark. This is achieved by designing a lightweight watermark embedding framework that operates in the common reverse diffusion paradigm shared by all TTS diffusion models. To mitigate the impact on audio quality, Smark utilizes the discrete wavelet transform (DWT) to embed watermarks into the relatively stable low-frequency regions of the audio, which ensures seamless watermark-audio integration and is resistant to removal during the reverse diffusion process. Extensive experiments are conducted to evaluate the audio quality and watermark performance in various simulated real-world attack scenarios. The experimental results show that Smark achieves superior performance in both audio quality and watermark extraction accuracy.

语音生成水印技术扩散模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。