arXiv:2504.10782cs.SDeess.AS2025-04被引 18

现有语音水印易被攻击移除,且无需知道水印方法

Deep Audio Watermarks are Shallow: Limitations of Post-Hoc Watermarking Techniques for Speech

  • 后处理方式嵌入水印,仅修改音频低层特征
  • 压缩过滤等操作可无损移除水印,音质损失小
  • 适用于评估语音水印安全性的研究者

在音频领域,最先进的水印技术利用深度神经网络在生成音频中嵌入人耳不可察觉的标识。理想情况是,即使音频经过压缩、滤波等变换,仍能高精度检测出水印。现有方法采用事后处理策略,在生成后通过添加微弱水印信号来操纵音频的“低层”特征。本文表明,这种事后设计使水印容易受到基于变换的移除攻击。针对语音音频,我们(1)统一并扩展了现有对音频变换影响水印可检测性的评估;(2)证明当前最先进的事后水印可在不了解水印方案的前提下被移除,且音频质量下降极小。

原文摘要 · Abstract (English)

In the audio modality, state-of-the-art watermarking methods leverage deep neural networks to allow the embedding of human-imperceptible signatures in generated audio. The ideal is to embed signatures that can be detected with high accuracy when the watermarked audio is altered via compression, filtering, or other transformations. Existing audio watermarking techniques operate in a post-hoc manner, manipulating "low-level" features of audio recordings after generation (e.g. through the addition of a low-magnitude watermark signal). We show that this post-hoc formulation makes existing audio watermarks vulnerable to transformation-based removal attacks. Focusing on speech audio, we (1) unify and extend existing evaluations of the effect of audio transformations on watermark detectability, and (2) demonstrate that state-of-the-art post-hoc audio watermarks can be removed with no knowledge of the watermarking scheme and minimal degradation in audio quality.

音频水印安全评估语音生成

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。