将水印嵌入语音离散表示,提升抗攻击能力与隐蔽性
Speech Watermarking with Discrete Intermediate Representations
- 通过向语音的离散潜在空间注入水印,增强鲁棒性
- 1秒语音可嵌入1至150比特水印,性能领先
- 适合语音克隆检测与信息隐藏场景
语音水印技术可主动缓解即时语音克隆带来的潜在危害。这类技术在语音中插入人类难以察觉但算法可识别的信号。以往方法通常在连续空间嵌入水印,而本文提出一种新框架DiscreteWM,将水印注入语音的离散中间表示。具体而言,利用向量量化自编码器将语音映射到离散潜在空间,并通过改变离散ID的模运算关系实现水印注入。为保证水印不可感知,还设计了一个操控模型以选择候选标记进行嵌入。实验表明,该框架在鲁棒性和不可感知性上均达到当前最优。此外,其灵活的帧级策略适用于语音克隆检测与信息隐藏。单段1秒语音可编码1至150比特水印信息。音频样例见https://DiscreteWM.github.io/discrete_wm。
原文摘要 · Abstract (English)
Speech watermarking techniques can proactively mitigate the potential harmful consequences of instant voice cloning techniques. These techniques involve the insertion of signals into speech that are imperceptible to humans but can be detected by algorithms. Previous approaches typically embed watermark messages into continuous space. However, intuitively, embedding watermark information into robust discrete latent space can significantly improve the robustness of watermarking systems. In this paper, we propose DiscreteWM, a novel speech watermarking framework that injects watermarks into the discrete intermediate representations of speech. Specifically, we map speech into discrete latent space with a vector-quantized autoencoder and inject watermarks by changing the modular arithmetic relation of discrete IDs. To ensure the imperceptibility of watermarks, we also propose a manipulator model to select the candidate tokens for watermark embedding. Experimental results demonstrate that our framework achieves state-of-the-art performance in robustness and imperceptibility, simultaneously. Moreover, our flexible frame-wise approach can serve as an efficient solution for both voice cloning detection and information hiding. Additionally, DiscreteWM can encode 1 to 150 bits of watermark information within a 1-second speech clip, indicating its encoding capacity. Audio samples are available at https://DiscreteWM.github.io/discrete_wm.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。