arXiv:2608.22186cs.SDcs.CL2026-08

无需训练即可为语音生成模型加水印,还能抵抗强干扰。

AudioNoisePrints: Model-free audio watermarking using spatial correlation in flow matching TTS

论文配图:AudioNoisePrints: Model-free audio watermarking using spatial correlation in flow matching TTS
图 1 · 摘自论文原文
  • 利用初始噪声与生成语音的强空间相关性实现水印嵌入。
  • 在F5TTS等模型上抗干扰能力优于AudioSeal基线。
  • 轻量检测器可提升水印强度,适合语音生成领域应用。

我们提出AudioNoisePrints,一种针对流匹配和扩散型文本转语音(TTS)模型的无训练水印方案,推理时额外计算极少,且不需重新训练模型或降低生成质量。该方法利用扩散与流匹配模型中初始高斯噪声与生成输出间的强空间相关性,通过简单的余弦相关性即可实现水印嵌入。此外,我们在其上训练了一个轻量级检测器以支持更激进的增强操作。实验表明,该方法在F5TTS及其他TTS与声码器模型上均表现优异,显著超越AudioSeal这一强基线。结果表明,这些模型普遍具备类似的空间相关性特征,预示该水印方案未来可推广至更多流匹配类TTS模型甚至声码器。

原文摘要 · Abstract (English)

We present AudioNoisePrints, a training-free watermarking pipeline for flow matching and diffusion TTS models, which requires minimal extra computation during inference and does not require retraining the TTS model or reducing the generation quality. We exploited the fact that there are strong correlations between the initial Gaussian noises and the generated outputs in diffusion and flow matching models, such that a simple cosine correlation between the initial noise and the generated output can be used to perform watermaking. Moreover, we train a lightweight detector on top for more aggressive augmentations. Our method outperforms AudioSeal, a strong baseline for audio watermarking under strong augmentations. We experimented on F5TTS and other TTS and vocoder models, and concluded that they all exhibit similar spatial correlation properties, suggesting our watermarking scheme can be used for more flow-matching TTS models and even vocoders in the future.

语音生成水印扩散模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。