用深度学习在音频中嵌入可验证水印,区分真假语音。
StreamMark: A Deep Learning-Based Semi-Fragile Audio Watermarking for Proactive Deepfake Detection

- 通过编码-失真-解码结构,利用复数域嵌入实现语义感知水印。
- 水印对压缩等正常处理鲁棒,对换声、编辑等攻击敏感(恢复率降至50%)。
- 适合需要主动识别深度伪造音频的场景,如语音可信验证。
生成式AI的快速发展使区分深度伪造音频与真实人声变得愈发困难。为克服被动检测方法的局限性,本文提出StreamMark,一种基于深度学习的半脆弱音频水印系统。该系统设计为对保留语义的良性音频转换(如压缩、加噪)具有鲁棒性,同时对恶意的语义篡改操作(如语音转换、语音编辑)保持脆弱性。方法采用独特的编码-失真-解码架构,并引入复数域嵌入技术,显式训练以区分两类变换。全面基准测试表明,StreamMark具备高不可察觉性(信噪比24.16 dB,PESQ 4.20),对真实世界失真(如Opus编码)具有鲁棒性,且在一系列深度伪造攻击下表现出合理脆弱性——消息恢复准确率降至随机水平(约50%),同时对良性基于AI的风格迁移仍保持高鲁棒性(准确率>98%)。
原文摘要 · Abstract (English)
The rapid advancement of generative AI has made it increasingly challenging to distinguish between deepfake audio and authentic human speech. To overcome the limitations of passive detection methods, we propose StreamMark, a novel deep learning-based, semi-fragile audio watermarking system. StreamMark is designed to be robust against benign audio conversions that preserve semantic meaning (e.g., compression, noise) while remaining fragile to malicious, semantics-altering manipulations (e.g., voice conversion, speech editing). Our method introduces a complex-domain embedding technique within a unique Encoder-Distortion-Decoder architecture, trained explicitly to differentiate between these two classes of transformations. Comprehensive benchmarks demonstrate that StreamMark achieves high imperceptibility (SNR 24.16 dB, PESQ 4.20), is resilient to real-world distortions like Opus encoding, and exhibits principled fragility against a suite of deepfake attacks, with message recovery accuracy dropping to chance levels (~50%), while remaining robust to benign AI-based style transfers (ACC >98%).
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。