为自回归音频生成模型设计无失真水印,防伪造更可靠
Robust Distortion-Free Watermark for Autoregressive Audio Generation Models
- 用聚类思想让同簇词元等价,解决重分词不一致问题
- 在多个主流平台测试,水印可检测性显著优于现有方法
- 适合需要内容可信度的语音合成、防诈骗场景
自回归语音模型的快速发展推动了对话交互的进步,但也带来了冒名顶替、误导性录音等滥用风险。为保障数字媒体真实性,水印技术成为关键安全措施。传统基于统计的水印方法在自回归音频模型上面临「重分词不一致」问题——原始与重分词后的离散音频词元序列存在差异。为此,我们提出Aligned-IS,一种专为音频生成模型设计的无失真水印方法。该方法通过聚类处理,使同一簇内的词元被视为等价,有效缓解重分词不一致带来的干扰。在多个主流音频生成平台上的全面测试表明,Aligned-IS不仅保持生成音频质量,还显著提升水印可检测性,超越当前最先进的无失真水印方案,确立了安全音频应用的新基准。
原文摘要 · Abstract (English)
The rapid advancement of next-token-prediction models has led to widespread adoption across modalities, enabling the creation of realistic synthetic media. In the audio domain, while autoregressive speech models have propelled conversational interactions forward, the potential for misuse, such as impersonation in phishing schemes or crafting misleading speech recordings, has also increased. Security measures such as watermarking have thus become essential to ensuring the authenticity of digital media. Traditional statistical watermarking methods used for autoregressive language models face challenges when applied to autoregressive audio models, due to the inevitable ``retokenization mismatch'' - the discrepancy between original and retokenized discrete audio token sequences. To address this, we introduce Aligned-IS, a novel, distortion-free watermark, specifically crafted for audio generation models. This technique utilizes a clustering approach that treats tokens within the same cluster equivalently, effectively countering the retokenization mismatch issue. Our comprehensive testing on prevalent audio generation platforms demonstrates that Aligned-IS not only preserves the quality of generated audio but also significantly improves the watermark detectability compared to the state-of-the-art distortion-free watermarking adaptations, establishing a new benchmark in secure audio technology applications.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。