arXiv:2609.03107cs.SDcs.CR2026-09

提出抗神经编码器的音频水印框架,保障合成语音可溯源性。

CRAW: Codec Robust Audio Watermarking

论文配图:CRAW: Codec Robust Audio Watermarking
图 1 · 摘自论文原文
  • 联合畸变感知训练与注意力池化提升鲁棒性
  • 在多种神经编码器下仍保持高检测率(>95%)
  • 适合用于真实场景中生成语音的版权保护

生成式语音模型使真实与合成音频难以区分,催生新型欺诈与虚假信息。音频水印通过嵌入不可察觉信号,可在事后验证语音来源。然而,现有事后水印方法在神经编码器和降噪器作用下失效,而这些变换常用于实际存储、传输与处理,严重限制其应用。本文提出CRAW框架,同时增强对神经重合成的鲁棒性并保持高感知质量。CRAW结合畸变感知训练、基于注意力的池化机制、推理时感知掩码及纠错码,以恢复鲁棒训练中损失的保真度。实验表明,CRAW在神经编码器、降噪器和声码器下均达到当前最优鲁棒性,且感知质量与现有事后水印方法相当。代码已开源:https://github.com/DavidC1212/craw。

原文摘要 · Abstract (English)

Recent advances in generative speech models have made it increasingly difficult to distinguish authentic from synthetic audio, enabling new forms of fraud and misinformation. Audio watermarking offers a promising defense by embedding an imperceptible signal into generated speech that can later be detected to verify its provenance. However, recent studies have shown that existing post-hoc watermarking methods fail under neural codecs and denoisers, transformations routinely applied during real-world storage, transmission, and processing, severely limiting their practical utility. Here we introduce CRAW, a codec-robust audio watermarking framework that jointly improves robustness against neural re-synthesis while maintaining high perceptual quality. CRAW combines distortion-aware training with an attention-based pooling mechanism, inference-time perceptual mask- ing, and an error-correcting code to recover the fidelity lost during robust training. Experiments demonstrate that CRAW achieves state-of-the-art robustness against neural codecs, denoisers, and vocoders while maintaining perceptual quality comparable to existing post-hoc watermarking methods. The code is available at https://github.com/DavidC1212/craw.

音频水印生成语音鲁棒性神经编码器

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。