arXiv:2606.21365cs.SDcs.AI2026-06

首次实现语义级音频水印,防篡改且能被克隆模型学习。

LambdaMark: Semantic Audio Watermarking for Robustness and Radioactivity

论文配图:LambdaMark: Semantic Audio Watermarking for Robustness and Radioactivity
图 1 · 摘自论文原文
  • 将多比特水印嵌入语义音频特征,而非波形或频谱
  • 在常见失真下水印恢复率接近100%,对移除攻击鲁棒
  • 适合用于语音克隆防护,尤其适用于生成模型输出

生成式音频技术的发展使得语音克隆变得极易实现,导致语音欺诈、冒用等未经授权的使用问题。常见攻击通过在目标说话人录音上微调语音生成模型,使其能够合成该说话人的声音。音频水印提供了一种有前景的防御手段,通过在音频中嵌入可检测信号来追踪来源。一个实用的水印需具备两个关键特性:鲁棒性和放射性。现有方法通常将信号嵌入低层表示(如波形或频谱),易受信号级操作影响,且难以传递到下游模型。我们提出LambdaMark——首个通用放射性水印方案。不同于以往方法,LambdaMark通过将多比特水印信息嵌入语义音频潜在表示中实现通用放射性。水印具有语义可解释性,因此更可能通过微调被下游模型学习。LambdaMark包含一个轻量级水印编码器,用于注入与消息相关的扰动,并配备解码器以检测水印并恢复比特信息。编码器与解码器通过定制的多组件损失函数联合训练,兼顾音频保真度、比特恢复率及对常见失真和对抗性移除攻击的鲁棒性。实验表明,LambdaMark在常见失真下达到近乎完美的鲁棒性,是唯一对所有评估过的移除攻击均保持鲁棒的水印方案。此外,其具备普遍且稳健的放射性,在微调模型生成的输出上仍能抵抗失真和对抗性移除攻击。

原文摘要 · Abstract (English)

Recent advances in generative audio have made voice cloning increasingly effortless, enabling voice fraud, impersonation, and other forms of unauthorized use. A common attack finetunes a speech generation model on recordings of a target speaker, allowing the model to synthesize speech in that speaker's voice. Audio watermarking offers a promising defense by embedding detectable signals into audio. A practical watermark must satisfy two key properties: robustness and radioactivity. Existing audio watermarking methods typically embed signals into low-level representations, such as waveforms or spectrograms, which makes them vulnerable to signal-level manipulations and limits their transfer to downstream models. We introduce LambdaMark -- the first generic radioactive watermarking scheme. Unlike all previous approaches, LambdaMark achieves generic radioactivity by embedding multi-bit watermark information into semantic audio latent representations. Our watermarks have semantic interpretation and are thus more likely to be learned by a downstream model through finetuning. LambdaMark includes a lightweight watermark encoder to inject multi-bit message-dependent perturbations into semantic audio representations and a decoder to detect watermark presence and recover the embedded bit information. Encoder and decoder are trained using a custom multi-component loss that preserves fidelity of the watermarked audio, increases bit-level recovery rate, and improves robustness against common distortions and adversarial removal attempts. Experiments show that LambdaMark achieves near-perfect robustness under common distortions. LambdaMark is also the only watermark that is robust against all evaluated removal attacks. Furthermore, LambdaMark exhibits general and robust radioactivity and remains robust to distortions and adversarial removal attacks even on the generated outputs of those finetuned models.

音频水印语音克隆鲁棒性放射性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。