用自嵌入隐写术实现无需训练的语音深度伪造主动防御
A Training-Free Proactive Defense Against Partial Speech Manipulation via Self-Embedding Steganography

- 通过自嵌入方式将语音自我压缩信息隐藏其中
- 在部分伪造场景下仍能准确检测并恢复原始语音
- 无需训练,适合资源受限的实时防御场景
部分深度伪造语音指仅对语句中有限片段进行合成或篡改,这对现有检测系统构成严峻挑战。随着伪造区域比例下降,被动检测方法可靠性显著降低,准确检测与修复难度增大。本文从新视角重新审视音频隐写术,提出将其作为对抗部分深度伪造的主动防御手段。具体而言,采用自嵌入策略,将干净语音信号自身压缩表示嵌入其中,实现事后参考内容提取。我们展示了现有音频隐写方法如何通过编解码器还原支持部分深度伪造检测。在基准数据集上的实验表明,该方法可有效补充被动防御。尤为关键的是,该方法无需任何训练,提供了一种鲁棒且数据高效的局部深度伪造检测方案。
原文摘要 · Abstract (English)
Partial deepfake speech, where only limited segments of an utterance are synthesized or manipulated, poses a significant challenge to existing deepfake detection systems. As the proportion of spoofed regions decreases, passive detectors become increasingly unreliable, and accurate detection and restoration remain challenging. In this paper, we revisit audio steganography from a new perspective and propose its use as a proactive defense against partially deepfaked audio. In particular, we consider a self-embedding strategy in which a clean speech signal embeds a compressed representation of itself, enabling post-hoc extraction of reference content. We demonstrate how existing audio steganography methods can be repurposed to support detection of partial deepfakes through codec-based restoration. Experiments on a benchmark dataset show that the proposed approach complements passive defenses. Remarkably, the proposed method operates without any training, providing a robust and data-efficient alternative for partial deepfake detection.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。