arXiv:2509.19812cs.SDcs.MM2025-09中稿 · ASRU 2025被引 1

用渐进式知识蒸馏实现高效抗攻击语音水印。

Efficient Speech Watermarking for Speech Synthesis via Progressive Knowledge Distillation

  • 用教师-学生结构蒸馏高保真语音水印模型
  • 计算成本降低93.6%,检测F1达99.6%
  • 适合实时语音合成场景的轻量级水印

随着语音生成模型快速发展,未经授权的语音克隆带来严重隐私与安全风险。语音水印为溯源和防滥用提供可行方案。现有技术主要分为基于信号处理(DSP)和深度学习两类:前者效率高但易受攻击,后者防护强但计算开销大。为此,我们提出PKDMark,一种基于渐进式知识蒸馏(PKD)的轻量级深度学习语音水印方法。该方法分两阶段进行:(1) 使用可逆神经网络架构训练高性能教师模型;(2) 通过渐进知识蒸馏将教师能力迁移至紧凑的学生模型。该过程使计算成本降低93.6%,同时保持高鲁棒性与不可察觉性。实验表明,蒸馏后模型在复杂失真下平均检测F1分数达99.6%,PESQ为4.30,适用于实时语音合成场景。

原文摘要 · Abstract (English)

With the rapid advancement of speech generative models, unauthorized voice cloning poses significant privacy and security risks. Speech watermarking offers a viable solution for tracing sources and preventing misuse. Current watermarking technologies fall mainly into two categories: DSP-based methods and deep learning-based methods. DSP-based methods are efficient but vulnerable to attacks, whereas deep learning-based methods offer robust protection at the expense of significantly higher computational cost. To improve the computational efficiency and enhance the robustness, we propose PKDMark, a lightweight deep learning-based speech watermarking method that leverages progressive knowledge distillation (PKD). Our approach proceeds in two stages: (1) training a high-performance teacher model using an invertible neural network-based architecture, and (2) transferring the teacher's capabilities to a compact student model through progressive knowledge distillation. This process reduces computational costs by 93.6% while maintaining high level of robust performance and imperceptibility. Experimental results demonstrate that our distilled model achieves an average detection F1 score of 99.6% with a PESQ of 4.30 in advanced distortions, enabling efficient speech watermarking for real-time speech synthesis applications.

语音水印知识蒸馏轻量化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。