arXiv:2511.02278eess.AS2025-11

用多重水印技术提升语音防伪抗攻击能力

Multiplexing Neural Audio Watermarks

  • 融合多种水印方法,通过时频域自适应设计增强鲁棒性
  • 在14种攻击下性能显著优于单一水印方案,抗干扰能力更强
  • 适合语音版权保护与真实场景中的防伪造应用

语音水印对验证语音真实性至关重要,但传统单水印方案在神经重建和对抗攻击等复杂干扰下表现不佳。为此,本文提出多路复用范式,结合多种水印技术以利用其互补优势。研究了并行与串行两种复用策略,提出感知自适应时频复用(PA-TFM),一种无需训练的鲁棒方法。为进一步提升性能,引入基于模型的MaskNet框架,用于学习有效时域复用。在LibriSpeech与Common Voice数据集上,针对14种不同攻击类型(包括高强度白盒攻击与神经重建攻击)的实验表明,PA-TFM与MaskNet均显著优于现有单水印基线,建立起适用于真实场景的强健语音保护范式。

原文摘要 · Abstract (English)

Audio watermarking is essential for verifying speech authenticity, yet single-watermark schemes often struggle against sophisticated distortions such as neural reconstruction and adversarial attacks. To address this limitation, we introduce a multiplexing paradigm that combines multiple watermarking techniques to leverage their inherent complementarities. We explore both parallel and sequential multiplexing strategies and propose perceptual-adaptive time-frequency multiplexing (PA-TFM), a robust training-free approach. To further enhance performance, we introduce MaskNet, a novel model-based framework designed to learn effective time-domain multiplexing. Experimental results on the LibriSpeech and Common Voice datasets under 14 diverse attack types, including high-strength white-box and neural reconstruction attacks, demonstrate that both PA-TFM and MaskNet considerably outperform existing single-watermark baselines, establishing a resilient paradigm for real-world audio protection.

音频水印语音安全对抗攻击多路复用

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。