首个端到端联合训练的语音编解码水印系统,提升真伪验证可靠性。
WMCodec: End-to-End Neural Speech Codec with Deep Watermarking for Authenticity Verification
- 端到端联合优化压缩与水印嵌入提取,提升隐秘性与可提取性。
- 在6kbps下16bps容量时,水印提取准确率超99%,抗攻击能力强。
- 适合需要高可信语音认证的场景,如语音通信、数字版权保护。
近期语音伪造技术的发展迫切需要更强的神经语音编解码器真实性验证机制。现有方法在压缩前嵌入数值水印,并从重建语音中提取验证,但存在水印与编解码分别训练、跨模态信息融合不足等问题,导致水印隐蔽性差、提取精度低、容量有限。为此,我们提出WMCodec,首个将压缩-重建与水印嵌入-提取端到端联合训练的神经语音编解码器,优化水印隐蔽性与可提取性。此外,设计迭代注意力印记单元(AIU),增强水印与语音特征的深层融合,降低量化噪声对水印的影响。实验表明,WMCodec在多数质量指标上优于AudioSeal with Encodec,且在水印提取准确性上持续领先AudioSeal with Encodec和增强版TraceableSpeech。在6 kbps带宽、16 bps水印容量下,面对常见攻击仍保持超过99%的提取准确率,展现出强鲁棒性。
原文摘要 · Abstract (English)
Recent advances in speech spoofing necessitate stronger verification mechanisms in neural speech codecs to ensure authenticity. Current methods embed numerical watermarks before compression and extract them from reconstructed speech for verification, but face limitations such as separate training processes for the watermark and codec, and insufficient cross-modal information integration, leading to reduced watermark imperceptibility, extraction accuracy, and capacity. To address these issues, we propose WMCodec, the first neural speech codec to jointly train compression-reconstruction and watermark embedding-extraction in an end-to-end manner, optimizing both imperceptibility and extractability of the watermark. Furthermore, We design an iterative Attention Imprint Unit (AIU) for deeper feature integration of watermark and speech, reducing the impact of quantization noise on the watermark. Experimental results show WMCodec outperforms AudioSeal with Encodec in most quality metrics for watermark imperceptibility and consistently exceeds both AudioSeal with Encodec and reinforced TraceableSpeech in extraction accuracy of watermark. At bandwidth of 6 kbps with a watermark capacity of 16 bps, WMCodec maintains over 99% extraction accuracy under common attacks, demonstrating strong robustness.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。