提出一种兼顾语音保真与鲁棒性的声纹水印方法
Protecting Your Voice: Temporal-aware Robust Watermarking
- 设计时序感知的轻量级编码器,提升水印重建质量
- 采用门控卷积网络实现逐位水印恢复,保持高鲁棒性
- 在真实语音和歌声上表现优异,平均PESQ达4.63
生成模型快速发展导致真假语音难以分辨。为消除模糊,常在频域特征中嵌入水印,但易损害语音细节保真度。本文提出一种时序感知的鲁棒水印方法(True),通过内容驱动的轻量级编码器实现波形重建,并设计时序感知门控卷积网络,逐位恢复水印信息。大量实验表明,该方法在保持高保真度的同时具备强鲁棒性,平均PESQ得分达4.63,优于现有最先进方法。
原文摘要 · Abstract (English)
The rapid advancement of generative models has led to the synthesis of real-fake ambiguous voices. To erase the ambiguity, embedding watermarks into the frequency-domain features of synthesized voices has become a common routine. However, the robustness achieved by choosing the frequency domain often comes at the expense of fine-grained voice features, leading to a loss of fidelity. Maximizing the comprehensive learning of time-domain features to enhance fidelity while maintaining robustness, we pioneer a \textbf{\underline{t}}emporal-aware \textbf{\underline{r}}ob\textbf{\underline{u}}st wat\textbf{\underline{e}}rmarking (\emph{True}) method for protecting the speech and singing voice. For this purpose, the integrated content-driven encoder is designed for watermarked waveform reconstruction, which is structurally lightweight. Additionally, the temporal-aware gated convolutional network is meticulously designed to bit-wise recover the watermark. Comprehensive experiments and comparisons with existing state-of-the-art methods have demonstrated the superior fidelity and vigorous robustness of the proposed \textit{True} achieving an average PESQ score of 4.63.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。