arXiv:2606.23335cs.SDcs.AI2026-06

水印反噬:语音水印让假声检测失效,反而把真声当假声。

The Watermark Shortcut: How Provenance Marking Sabotages Audio Deepfake Detection

  • 检测器误将水印当作假声标志,形成错误依赖。
  • 水印真声被误判为假声,等错误率飙升至75%。
  • 重训模型并加水印到真声数据可解决该问题。

语音溯源水印正被广泛用于合成语音防护,如Chatterbox模型内置、AudioSeal技术或ElevenLabs平台部署。我们发现一种未被识别的隐患:当合成语音有水印而真人语音无水印时,检测器会错误地将水印视为'假声'的捷径特征。这一错误依赖导致三大问题:泛化能力下降(在未见数据上表现变差)、去水印逃逸(水印去除后假声可逃脱检测)、水印误标(给真实语音加水印会使其被标记为假声)。在白盒实验中,水印训练的检测器出现全部三类失败(例如,等错误率从16%升至75%)。黑盒测试显示,向真实语音添加水印可使其被误判为假声。但该问题可修复:将水印同时应用于真声与假声数据再训练模型,即可消除水印依赖,恢复正常检测行为。我们公开了配对的纯净与带水印语音语料库(WASP)。

原文摘要 · Abstract (English)

Provenance watermarking is increasingly treated as a safeguard for synthetic speech, whether built directly into speech-generation models such as Chatterbox, provided through dedicated techniques such as AudioSeal, or deployed by commercial platforms such as ElevenLabs. We identify a previously uncharacterized liability: when synthetic speech is watermarked and human speech is not, detectors trained alongside latch onto the watermark as a spurious "watermark => fake" shortcut. This single feature yields three coupled failures: generalization degradation (model performance deteriorates on unseen data), strip-to-evade (a watermarked fake escapes once unwatermarked), and mark-to-frame (watermarking a real voice flags it as fake). In a controlled white-box experiment, a watermark-trained detector shows all three (for example, mark-to-frame lifts Equal Error Rate from 16% to 75%). In a black-box test of a commercial API, we show that adding a watermark to real speech disguises it as fake. However, this shortcut is fixable: retraining with the watermark on both classes decorrelates it and restores clean behavior. We release experiment data as a paired clean-versus-watermarked corpus (WASP).

语音生成水印检测失效深度伪造

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。