arXiv:2601.20432cs.SDcs.AI2026-01中稿 · IEICE, SP/SLP 2026被引 4

用自语音转换攻击音频水印,能破坏现有技术可靠性。

Self Voice Conversion as an Attack against Neural Audio Watermarking

  • 用语音转换模型保持内容但改变音色,实现无损攻击。
  • 实验证明该攻击使主流水印方法失效,误检率显著上升。
  • 适用于评估水印系统安全性的研究人员与开发者。

音频水印在保持说话人身份、语言内容和听觉质量的前提下嵌入辅助信息。尽管基于神经网络和数字信号处理的水印方法在不可感知性和嵌入容量方面取得进展,其鲁棒性仍主要针对压缩、加性噪声和重采样等传统失真进行评估。然而,深度学习带来的新型攻击对水印安全性构成重大威胁。本文研究自语音转换作为一种通用、内容保持型攻击对音频水印系统的影响。自语音转换通过语音转换模型将说话人音色映射为同一身份,同时改变声学特征。我们证明该攻击严重削弱了先进水印方法的可靠性,并揭示其对现代音频水印技术安全性的深远影响。

原文摘要 · Abstract (English)

Audio watermarking embeds auxiliary information into speech while maintaining speaker identity, linguistic content, and perceptual quality. Although recent advances in neural and digital signal processing-based watermarking methods have improved imperceptibility and embedding capacity, robustness is still primarily assessed against conventional distortions such as compression, additive noise, and resampling. However, the rise of deep learning-based attacks introduces novel and significant threats to watermark security. In this work, we investigate self voice conversion as a universal, content-preserving attack against audio watermarking systems. Self voice conversion remaps a speaker's voice to the same identity while altering acoustic characteristics through a voice conversion model. We demonstrate that this attack severely degrades the reliability of state-of-the-art watermarking approaches and highlight its implications for the security of modern audio watermarking techniques.

音频水印语音转换安全攻击

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。