arXiv:2606.15264eess.AScs.SD2026-06中稿 · INTERSPEECH 2026

用音节时长编辑实现抗生成攻击的语音水印。

DuraMark: Duration-Embedded Watermarking in LLM-based TTS

  • 通过控制音节时长嵌入水印信息。
  • 在多种生成攻击下仍保持90%以上检测率。
  • 适合语音版权保护与深度伪造防范场景。

基于大语言模型(LLM)的文本转语音(TTS)模型具备强大的语音克隆能力,引发深度伪造滥用风险。语音水印通过在生成语音中嵌入可追溯信息以缓解此问题。主流水印方法在波形或频谱层面操作,易受神经编解码器和声码器等生成攻击影响。为此,我们提出DuraMark,一种鲁棒的信息级水印框架,利用音节时长编辑实现水印嵌入。具体而言,DuraMark结合可控时长的LLM-TTS模型在合成过程中编辑音节时长,并配备时长提取器用于检测。实验表明,DuraMark在对抗生成攻击方面显著优于信号级基线方法。音频样例可访问 https://muzw.github.io/duramark_demo/。

原文摘要 · Abstract (English)

Large language model (LLM)-based text-to-speech (TTS) models have achieved remarkable voice cloning capabilities, raising concerns about potential deepfake misuse. Speech watermarking mitigates this by embedding traceable information into generated speech. Mainstream watermarking methods operate at the signal level (waveform or spectrogram), rendering the watermark vulnerable to generative attacks (e.g., neural codec and vocoder). To address this, we propose DuraMark, a robust information-level watermarking framework. It utilizes syllable duration editing to achieve watermark embedding. Specifically, DuraMark integrates a duration-controllable LLM-based TTS model to edit syllable durations during synthesis, coupled with a duration extractor to extract these durations for detection. Experiments demonstrate DuraMark's superior robustness against generative attacks, significantly outperforming signal-level baselines. Audio samples are available at https://muzw.github.io/duramark_demo/.

语音水印语音克隆生成攻击时长编辑

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。