arXiv:2608.24159cs.SD2026-08

音频水印会严重削弱部分语音深伪检测系统性能,暴露其鲁棒性漏洞。

On the Robustness of Audio Deepfake Detection under Audio Watermarking

论文配图:On the Robustness of Audio Deepfake Detection under Audio Watermarking
图 1 · 摘自论文原文
  • 将音频水印视为非对抗性扰动,评估其对多种检测模型的影响。
  • 在ASVspoof 2021 LA和DF数据集上检测准确率显著下降,最高降幅达30%。
  • 水印导致嵌入空间特征偏移,是检测失效的关键原因,适合安全与可信语音研究者参考。

生成式语音模型的进步使得合成语音高度逼真,因此可靠的声音深伪检测(ADD)系统变得愈发重要。尽管先前研究主要关注对抗性扰动,但对真实信号变换下ADD系统的鲁棒性仍缺乏充分理解。本文将音频水印视为结构化、非对抗性扰动,而非传统攻击手段,基于WavMark构建评估框架,测试了多种自监督学习(SSL)、卷积神经网络(CNN)及图神经网络(GNN)基的ADD模型在多个基准数据集上的表现。除了常规检测指标外,还通过弗雷歇距离、余弦相似度和L2距离分析水印引起的嵌入空间表示变化。实验结果表明,水印影响具有显著数据集依赖性:在ASVspoof 2021 LA和DF上造成显著性能下降,而在ASVspoof 2024、FoR和ITW上影响较小。此外,大范围的嵌入空间偏移与检测性能急剧恶化密切相关,说明水印可显著改变当前检测系统所依赖的特征表示。这表明,为内容保护设计的良性信号变换可能暴露出语音深伪检测系统此前未被注意的鲁棒性缺陷。代码已开源于https://github.com/ziqian0925/wm-ADD-robustness.git。

原文摘要 · Abstract (English)

Recent advances in generative audio models have enabled highly realistic synthetic speech, increasing the importance of reliable audio deepfake detection (ADD) systems. While prior studies have primarily focused on adversarially optimized perturbations, the robustness of ADD systems under realistic signal transformations remains insufficiently understood. In this work, we investigate the impact of audio watermarking on ADD systems by treating watermarking as a structured, non-adversarial perturbation rather than a conventional attack mechanism. Using a watermark-based evaluation framework built upon WavMark, we evaluate multiple self-supervised learning (SSL), Convolutional Neural Network (CNN) and Graph Neural Netrowk (GNN)-based ADD models across several benchmark datasets. Beyond conventional detection metrics, we further analyze watermark-induced representation shifts using Fréchet Distance, cosine similarity, and L2 distance in the embedding space. Experimental results reveal a strong dataset-dependent behavior: watermarking causes substantial performance degradation on ASVspoof 2021 LA and DF, while exhibiting limited impact on ASVspoof 2024, FoR, and ITW. Moreover, large embedding-space shifts are strongly associated with severe detection degradation, suggesting that watermark-induced perturbations can substantially alter the feature representations relied upon by current ADD systems. These findings demonstrate that benign signal transformations designed for content protection can expose previously overlooked robustness vulnerabilities in audio deepfake detection systems. Our code is available at https://github.com/ziqian0925/wm-ADD-robustness.git

音频安全深伪检测水印攻击鲁棒性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。