arXiv:2607.21132cs.SD2026-07

在音频编码器潜空间嵌入水印,提升对神经编码的鲁棒性。

Investigating Codec-Internal Latent Audio Watermarking for Neural Codec Robustness

论文配图:Investigating Codec-Internal Latent Audio Watermarking for Neural Codec Robustness
图 1 · 摘自论文原文
  • 将32位信息嵌入编码器的连续潜变量中,而非波形或频谱
  • 在48kHz语音上,水印准确率从78.8%提升至95.6%~97.1%
  • 兼顾抗神经编码破坏能力,适合语音版权保护场景

神经音频编码器对音频水印构成挑战,因其会重新编码、量化并重合成语音。本文研究了针对编码器鲁棒性的连续潜空间水印方法。不同于仅在波形或频谱上加水印,我们将在类似语音自编码器的编码器连续潜表示中嵌入32位消息。该流程采用SEANet风格的编解码器、基于Conformer的消息嵌入器、RVQ引导的潜变量分解,以及在信号处理和神经编码器变换下训练的潜域检测器。本文未提出通用基准,而是分析当水印载体移至神经解码前时出现的权衡。在48 kHz语音上,针对EnCodec的训练使EnCodec-24k比特下的水印准确率从78.8%提升至95.6%和97.1%,而PESQ值从3.727下降至3.514和3.427。

原文摘要 · Abstract (English)

Neural audio codecs are challenging transformations for audio watermarking because they re-encode, quantize, and resynthesize speech. This paper investigates continuous latent-space watermarking for codec robustness. Instead of adding a watermark only to the waveform or spectrogram, we embed a 32-bit message into the continuous latent representation of a codec-like speech autoencoder. The pipeline uses a SEANet-style encoder-decoder, a Conformer-based message embedder, RVQ-guided latent decomposition, and a latent-domain detector trained under signal-processing and neural-codec transformations. Rather than proposing a final universal watermarking baseline, we characterize the trade-offs that appear when the watermark carrier is moved before neural decoding. On 48 kHz speech, EnCodec-aware training improves EnCodec-24k bit accuracy from 78.8% to 95.6% and 97.1%, while PESQ decreases from 3.727 to 3.514 and 3.427.

音频水印潜空间编码鲁棒性语音安全

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。