提出无水印的可追溯语音合成框架,兼顾音质与溯源能力。
Traceable TTS: Toward Watermark-Free TTS with Strong Traceability
- 采用联合训练机制,不依赖水印实现语音溯源。
- 音质保持甚至略有提升,且溯源泛化性能更强。
- 适合关注语音安全与模型可追溯性的研究者。
近期文本转语音(TTS)技术的发展使合成语音在逼真度上接近人类发音,引发重大安全隐忧。因此亟需具备强溯源能力的可追溯TTS系统,同时不牺牲音质或安全性。然而现有方法多依赖语音或声码器中的显式水印,导致音质下降且易被伪造。为此,本文提出一种新型模型归属框架:不嵌入水印,而是通过联合训练优化TTS模型与判别器,显著提升溯源泛化能力,同时保持甚至略微改善音频质量。这是首个实现无水印、强溯源的TTS方案。为推动领域发展,论文录用后将开源代码。
原文摘要 · Abstract (English)
Recent advances in Text-To-Speech (TTS) technology have enabled synthetic speech to mimic human voices with remarkable realism, raising significant security concerns. This underscores the need for traceable TTS models-systems capable of tracing their synthesized speech without compromising quality or security. However, existing methods predominantly rely on explicit watermarking on speech or on vocoder, which degrades speech quality and is vulnerable to spoofing. To address these limitations, we propose a novel framework for model attribution. Instead of embedding watermarks, we train the TTS model and discriminator using a joint training method that significantly improves traceability generalization while preserving-and even slightly improving-audio quality. This is the first work toward watermark-free TTS with strong traceability. To promote progress in related fields, we will release the code upon acceptance of the paper.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。