arXiv:2504.15035cs.CRcs.AI2025-04被引 1

用低秩适配实现高效语音生成水印,抗攻击能力强且计算开销小。

SOLIDO: A Robust Watermarking Method for Speech Synthesis via Low-Rank Adaptation

  • 通过低秩适配实现参数高效微调,融合水印编码与扩散模型输入
  • 支持变长语音输入,水印提取准确率达99.20%(单类攻击)和98.43%(复合攻击)
  • 在时间拉伸攻击下优于现有方法近23%,适合需版权保护的语音生成场景

语音生成模型的快速发展带来了模型侵权和内容滥用等安全问题。现有生成水印技术大多计算开销大、训练成本高,且对变长输入鲁棒性不足。为此,我们提出SOLIDO,一种基于低秩适配(LoRA)的语音扩散模型生成水印方法。水印编码器将水印信息转换为适配扩散模型的输入形式;为实现变长输入下的精确水印提取,设计了基于深度可分离卷积的水印解码器。同时提出轻量级语音驱动微调策略,通过LoRA降低计算开销。大量实验表明,该方法在2000 bps高容量下仍能生成高保真水印语音。面对常见单类与复合语音攻击,水印提取平均准确率分别达99.20%和98.43%;在时间拉伸攻击中性能优于现有方法近23%。

原文摘要 · Abstract (English)

The accelerated advancement of speech generative models has given rise to security issues, including model infringement and unauthorized abuse of content. Although existing generative watermarking techniques have proposed corresponding solutions, most methods require substantial computational overhead and training costs. In addition, some methods have limitations in robustness when handling variable-length inputs. To tackle these challenges, we propose \textsc{SOLIDO}, a novel generative watermarking method that integrates parameter-efficient fine-tuning with speech watermarking through low-rank adaptation (LoRA) for speech diffusion models. Concretely, the watermark encoder converts the watermark to align with the input of diffusion models. To achieve precise watermark extraction from variable-length inputs, the watermark decoder based on depthwise separable convolution is designed for watermark recovery. To further enhance speech generation performance and watermark extraction capability, we propose a speech-driven lightweight fine-tuning strategy, which reduces computational overhead through LoRA. Comprehensive experiments demonstrate that the proposed method ensures high-fidelity watermarked speech even at a large capacity of 2000 bps. Furthermore, against common individual and compound speech attacks, our SOLIDO achieves a maximum average extraction accuracy of 99.20\% and 98.43\%, respectively. It surpasses other state-of-the-art methods by nearly 23\% in resisting time-stretching attacks.

语音生成水印技术扩散模型低秩适配

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。