arXiv:2606.11828cs.SDcs.AI2026-06中稿 · ICME2026被引 1

让音频水印更抗重建干扰,同时保持听不出痕迹。

Feature-Aligned Speech Watermarking for Robustness to Reconstruction Distortions

论文配图:Feature-Aligned Speech Watermarking for Robustness to Reconstruction Distortions
图 1 · 摘自论文原文
  • 水印与原始语音特征对齐,提升能量使用效率
  • 在可见与不可见重建模型下鲁棒性显著增强
  • 适合需要高鲁棒性的音频版权保护场景

音频水印旨在嵌入可识别信息的同时保持不可感知性。现有方法采用高保真、低能量设计以维持听觉质量,但导致水印在语音重建模型压制下缺乏鲁棒性。由于现有设计中存在鲁棒性-保真度权衡,提高水印能量虽能增强鲁棒性,却降低保真度。为此,我们提出一种特征对齐的水印方法,将水印与原始语音特征分布对齐,使更高能量的水印在提升鲁棒性的同时仍保持不可感知。利用预训练语音编解码器生成伪语音水印,并将其融合至输入音频的频谱图中,通过语音活动检测(VAD)损失和感知损失引导水印嵌入有声区域。实验表明,本方法在保持与现有方法相当的不可感知性的同时,在已见与未见语音重建模型下均显著提升了鲁棒性。

原文摘要 · Abstract (English)

Audio watermarking aims to embed identifiable information into audio while remaining imperceptible. Existing methods adopt high-fidelity, low-energy designs to preserve perceptual quality, but the resulting watermarks lack robustness under suppression by speech reconstruction models. Improving robustness is challenging due to the inherent robustness-fidelity trade-off in existing designs, where increasing watermark energy improves robustness but reduces fidelity. To address this problem, we propose a feature-aligned watermarking method that aligns the watermark with the original speech feature distribution, allowing higher watermark energy to improve robustness while preserving imperceptibility. We use a pretrained speech codec to generate a pseudo-speech watermark and fuse it into the spectrogram of the input audio, with VAD loss and perceptual losses guiding embedding within voiced regions. Experiments show that our method maintains imperceptibility comparable to existing approaches while substantially improving robustness under both seen and unseen speech reconstruction models.

音频水印语音重建鲁棒性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。