融合波形与频谱特征,提升语音伪造检测的泛化能力。
Two Views, One Truth: Spectral and Self-Supervised Features Fusion for Robust Speech Deepfake Detection
- 结合自监督表征与手工频谱特征,互补增强
- 交叉注意力策略使错误率降低38%
- 适合应对未知伪造攻击的鲁棒检测场景
近期合成语音技术的进步使音频深度伪造愈发逼真,带来严重安全风险。现有检测方法依赖单一模态(原始波形嵌入或频谱特征),易受非伪造干扰且对已知伪造算法过拟合,泛化能力差。为此,本文研究融合自监督学习(SSL)表征与手工频谱描述符(MFCC、LFCC、CQCC)的混合框架。通过跨模态对齐与融合,捕捉单特征方法忽略的细微伪造痕迹。探索了简单拼接、交叉注意力、互交叉注意力及可学习门控机制等多种融合策略,以最优结合SSL特征与精细频谱线索。在四个公开基准上评估,所有融合变体均优于仅用SSL的基线,其中交叉注意力策略实现最佳泛化性能,等错误率(EER)相对降低38%。结果表明,联合建模波形与频谱视图可生成鲁棒、领域无关的音频深度伪造检测表示。
原文摘要 · Abstract (English)
Recent advances in synthetic speech have made audio deepfakes increasingly realistic, posing significant security risks. Existing detection methods that rely on a single modality, either raw waveform embeddings or spectral based features, are vulnerable to non spoof disturbances and often overfit to known forgery algorithms, resulting in poor generalization to unseen attacks. To address these shortcomings, we investigate hybrid fusion frameworks that integrate self supervised learning (SSL) based representations with handcrafted spectral descriptors (MFCC , LFCC, CQCC). By aligning and combining complementary information across modalities, these fusion approaches capture subtle artifacts that single feature approaches typically overlook. We explore several fusion strategies, including simple concatenation, cross attention, mutual cross attention, and a learnable gating mechanism, to optimally blend SSL features with fine grained spectral cues. We evaluate our approach on four challenging public benchmarks and report generalization performance. All fusion variants consistently outperform an SSL only baseline, with the cross attention strategy achieving the best generalization with a 38% relative reduction in equal error rate (EER). These results confirm that joint modeling of waveform and spectral views produces robust, domain agnostic representations for audio deepfake detection.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。