通过消除语音中的说话人特征,让检测模型更关注合成痕迹。
SNAP: Speaker Nulling for Artifact Projection in Speech Deepfake Detection
- 用正交投影方法剥离说话人信息,聚焦合成伪影
- 在多个数据集上实现当前最优检测准确率
- 适合需要跨说话人泛化的语音伪造检测场景
近年来文本转语音技术可生成几乎无法与真实语音区分的合成语音。尽管基于自监督学习的语音编码器在深度伪造检测中表现良好,但其泛化能力受限于未见说话人。我们定量分析发现,这些编码表示严重受说话人信息影响,导致检测器依赖说话人特有相关性而非合成伪影线索。我们称之为说话人纠缠现象。为缓解此问题,提出SNAP框架:估计说话人子空间并应用正交投影,抑制说话人相关成分,将合成伪影保留在残差特征中。通过减少说话人纠缠,SNAP促使检测器关注伪影模式,显著提升性能,达到当前最优水平。
原文摘要 · Abstract (English)
Recent advancements in text-to-speech technologies enable generating high-fidelity synthetic speech nearly indistinguishable from real human voices. While recent studies show the efficacy of self-supervised learning-based speech encoders for deepfake detection, these models struggle to generalize across unseen speakers. Our quantitative analysis suggests these encoder representations are substantially influenced by speaker information, causing detectors to exploit speaker-specific correlations rather than artifact-related cues. We call this phenomenon speaker entanglement. To mitigate this reliance, we introduce SNAP, a speaker-nulling framework. We estimate a speaker subspace and apply orthogonal projection to suppress speaker-dependent components, isolating synthesis artifacts within the residual features. By reducing speaker entanglement, SNAP encourages detectors to focus on artifact-related patterns, leading to state-of-the-art performance.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。