arXiv:2504.20532cs.MMcs.CR2025-04被引 1

为扩散模型生成语音设计三重溯源水印,防伪追踪更全面。

TriniMark: A Robust Generative Speech Watermarking Method for Trinity-Level Traceability

  • 轻量编码器嵌入水印比特,时序感知解码器可靠恢复。
  • 支持内容、模型、用户三重溯源,抗多种信号攻击。
  • 单模型可变水印,适合大规模用户追踪场景。

基于扩散的语音生成已实现极高的保真度,但滥用和未经授权传播的风险随之上升。现有水印方法多针对GAN架构,对扩散模型的水印研究仍不充分。以往工作主要关注内容级溯源,而对模型级和用户级归属的支持较弱。本文提出TriniMark,一种面向扩散模型的生成语音水印框架,实现三重溯源能力:即关联生成语音样本与嵌入的水印信息(内容级溯源)、源生成模型(模型级归属)以及请求生成的终端用户(用户级追踪)。TriniMark采用轻量编码器将水印比特嵌入时域语音特征并重建波形,使用时序感知门控卷积解码器实现可靠比特恢复。我们进一步引入波形引导微调策略,将水印能力迁移至扩散模型。最后通过可变水印训练,使单一训练模型在推理时可嵌入不同水印信息,实现可扩展的用户级追踪。在多个语音数据集上的实验表明,TriniMark在保持语音质量的同时,提升了对常见单次及复合信号处理攻击的鲁棒性,并支持大容量水印以实现大规模溯源。

原文摘要 · Abstract (English)

Diffusion-based speech generation has achieved remarkable fidelity, increasing the risk of misuse and unauthorized redistribution. However, most existing generative speech watermarking methods are developed for GAN-based pipelines, and watermarking for diffusion-based speech generation remains comparatively underexplored. In addition, prior work often focuses on content-level provenance, while support for model-level and user-level attribution is less mature. We propose \textbf{TriniMark}, a diffusion-based generative speech watermarking framework that targets trinity-level traceability, i.e., the ability to associate a generated speech sample with (i) the embedded watermark message (content-level provenance), (ii) the source generative model (model-level attribution), and (iii) the end user who requested generation (user-level traceability). TriniMark uses a lightweight encoder to embed watermark bits into time-domain speech features and reconstruct the waveform, and a temporal-aware gated convolutional decoder for reliable bit recovery. We further introduce a waveform-guided fine-tuning strategy to transfer watermarking capability into a diffusion model. Finally, we incorporate variable-watermark training so that a single trained model can embed different watermark messages at inference time, enabling scalable user-level traceability. Experiments on speech datasets indicate that TriniMark maintains speech quality while improving robustness to common single and compound signal-processing attacks, and it supports high-capacity watermarking for large-scale traceability.

语音生成水印技术扩散模型溯源追踪

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。