在语音生成模型权重中嵌入可追溯水印,保护开源模型版权。
P2Mark: Plug-and-play Parameter-level Watermarking for Neural Speech Generation
- 将水印信息嵌入预训练模型权重,通过轻量适配器实现
- 在两种主流波形解码器上达到顶尖水印提取准确率和不可感知性
- 适合需要开源模型版权保护的研究者与开发者使用
神经语音生成(NSG)作为人工智能生成内容的关键组件,已能生成高质量、高度逼真的语音,广泛应用于各类场景。然而,该技术的滥用可能威胁社会安全。音频水印可通过在生成音频中嵌入不可感知的标记,为安全使用提供可行方案。现有方法多作用于音频或特征层面,不适用于源代码与模型权重公开的场景。为此,我们提出一种即插即用的参数级水印方法(P2Mark),将水印嵌入释放的模型权重中,实现对开源场景下模型版权的主动追踪与保护。训练时,我们在预训练模型中引入轻量水印适配器,使水印信息通过适配器融合至模型参数。该设计支持发布前灵活修改水印,且发布后水印安全固化。同时,采用梯度正交投影优化策略,保障生成音频质量与水印保真度。在两类主流波形解码器(声码器与编解码器)上的实验表明,P2Mark在水印提取准确率、不可感知性和鲁棒性方面,达到在非开源白盒保护场景下不可适用的顶尖方法水平。
原文摘要 · Abstract (English)
Neural speech generation (NSG) has rapidly advanced as a key component of artificial intelligence-generated content, enabling the generation of high-quality, highly realistic speech for diverse applications. This development increases the risk of technique misuse and threatens social security. Audio watermarking can embed imperceptible marks into generated audio, providing a promising approach for secure NSG usage. However, current audio watermarking methods are mainly applied at the audio-level or feature-level, which are not suitable for open-sourced scenarios where source codes and model weights are released. To address this limitation, we propose a Plug-and-play Parameter-level WaterMarking (P2Mark) method for NSG. Specifically, we embed watermarks into the released model weights, offering a reliable solution for proactively tracing and protecting model copyrights in open-source scenarios. During training, we introduce a lightweight watermark adapter into the pre-trained model, allowing watermark information to be merged into the model via this adapter. This design ensures both the flexibility to modify the watermark before model release and the security of embedding the watermark within model parameters after model release. Meanwhile, we propose a gradient orthogonal projection optimization strategy to ensure the quality of the generated audio and the accuracy of watermark preservation. Experimental results on two mainstream waveform decoders in NSG (i.e., vocoder and codec) demonstrate that P2Mark achieves comparable performance to state-of-the-art audio watermarking methods that are not applicable to open-source white-box protection scenarios, in terms of watermark extraction accuracy, watermark imperceptibility, and robustness.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。