用语义水印让合成语音更难被篡改且可追踪
SSTMark: Robust Training-Free Semantic-Level Speech Watermarking
- 在文本层面嵌入水印,而非波形或频谱
- 在信号处理和压缩攻击下检测率提升4.6%至16.9%
- 无需训练,适合快速部署于各类语音生成模型
随着语音生成模型日益逼真且普及,合成语音的滥用、归属认定与治理问题愈发突出。水印技术为合成语音提供可追溯与可验证的解决方案。现有方法多在信号层(如波形或频谱)嵌入水印,但在强干扰下易被破坏,导致检测能力下降。本文提出SSTMark,一种无需训练的语义级语音水印框架,通过文本水印实现语义层面的水印嵌入与检测。实验表明,在AudioMarkBench测试中,SSTMark平均鲁棒性最强。在固定误报率为1%时,相比最先进基线,其在信号处理编辑和压缩编辑下的平均检测率分别提升4.6%和16.9%。
原文摘要 · Abstract (English)
As speech generation models become increasingly realistic and widely accessible, concerns about the misuse, attribution, and governance of synthetic speech continue to grow. Watermarking provides a practical way to make synthesized speech traceable and verifiable. Most existing speech watermarking methods embed watermark information into signal-level representations, such as waveforms or spectrograms. Under sufficiently strong distortions, the embedded watermark may be weakened or destroyed, leading to degraded detectability. In this paper, we propose SSTMark, a training-free speech watermarking framework that operates at the semantic level through text watermarking. Unlike conventional signal-level watermarking methods, SSTMark encodes watermark information into the semantic content conveyed by generated speech, and detects the watermark from the recovered linguistic content. Experiments on AudioMarkBench demonstrate that SSTMark exhibits the strongest average robustness. Compared with the state-of-the-art baselines at a fixed false positive rate of 1\%, SSTMark improves the average detection rate by 4.6\% and 16.9\% on signal-processing edits and compression edits, respectively.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。