arXiv:2605.01515cs.SDcs.CR2026-05中稿 · ACISP 2026

在语音生成时嵌入水印,防伪溯源且不影响听感。

MelShield: Robust Mel-Domain Audio Watermarking for Provenance Attribution of AI Generated Synthesized Speech

论文配图:MelShield: Robust Mel-Domain Audio Watermarking for Provenance Attribution of AI Generated Synthesized Speech
图 1 · 摘自论文原文
  • 生成过程中对梅尔频谱加密扰动,隐蔽嵌入二进制信息。
  • 压缩和噪声下仍保持近100%水印提取准确率。
  • 无需修改模型,适合多用户版权追踪,防恶意破解。

本文提出 MelShield,一种鲁棒、生成时、密钥控制的音频水印框架,用于保护人工智能合成语音的版权并实现可靠溯源。该方法在文本转语音(TTS)生成流程中,针对梅尔频谱条件化架构的中间声学表示进行操作,将短二进制数据通过低能量、密钥控制的扩频扰动嵌入到特定时间-频率区域。由于在声码器推理前完成水印嵌入,MelShield可无缝适配各类梅尔条件化TTS架构(如DiffWave、HiFi-GAN),无需修改或重训练。其多用户密钥设计支持用户专属溯源,密钥验证机制可有效防止非授权解码与大规模对抗分析。在DiffWave和HiFi-GAN上的大量实验表明,即便在压缩、加性噪声等信号失真条件下,仍能实现接近100%的水印提取准确率,同时保持高感知音频质量。

原文摘要 · Abstract (English)

In this paper, we propose MelShield, a robust, in-generation, keyed audio watermarking framework that embeds identifiable signals into AI-generated audio for copyright protection and reliable attribution. Specifically, MelShield operates in the Mel-spectrogram domain during the generation process, targeting intermediate acoustic representations in Mel-conditioned pipelines for text-to-speech (TTS) generation. The core idea is to treat the intermediate Mel-spectrogram as the host signal and embed a short binary payload via low-energy, keyed spread-spectrum perturbations distributed across carefully selected time-frequency regions prior to waveform synthesis. By performing watermarking before vocoder inference, MelShield remains plug-and-play for Mel-conditioned TTS architectures and does not require modification or retraining of the underlying TTS generation vocoder, such as DiffWave and HiFi-GAN. Moreover, the multi-user keyed construction enables scalable user-specific attribution, while the keyed verification mechanism limits unauthorized decoding, thereby reducing the risk of large-scale extractor probing and adversarial analysis. Extensive experiments on DiffWave and HiFi-GAN demonstrate that MelShield achieves reliable watermark extraction, approaching 100\% bit accuracy, even under signal distortions, e.g., compression and additive noise, while preserving high perceptual audio quality.

语音生成水印版权保护扩散模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。