针对混合音频中的语音伪造检测难题,提出新数据集与多流提示调优方法。
MixFake: Benchmarking and Enhancing Audio Deepfake Detection in Diverse Real-world Mixed Audio

- 通过多流提示注入,融合时域、频域和纹理特征提升检测能力
- 在前景语音检测中实现0.95%的等错误率,复杂背景提升7.72%准确率
- 适合关注真实场景下语音安全的研究者与应用开发者
语音深度伪造检测在纯净环境下已取得显著进展,但在真实世界复杂场景中,语音常与背景音乐或噪声混合,面临严峻挑战。现有先进方法依赖自监督学习(SSL)模型的语义特征,处理非语音或混合源音频时表现不佳。本文首次提出MixFake,一个大规模基准数据集,模拟多种声学环境,涵盖不同信噪比(SNR)及真假语音混合比例。为突破“语义中心”局限,我们提出多流提示调优框架,将信号级先验信息注入SSL骨干网络。通过基线、频率和纹理三路流的深度提示融合,有效捕捉音频伪造痕迹。实验表明,该方法显著优于现有基线,在前景语音检测中达到0.95%等错误率(EER),在复杂背景任务中实现7.72%的绝对性能提升。数据集与代码已开源。
原文摘要 · Abstract (English)
Speech deepfake detection has achieved remarkable success in clean environments but faces significant challenges in complex, real-world scenarios where speech is often mixed with background music or noise. Current state-of-the-art methods rely on semantic features from self-supervised learning (SSL) models, which often fail when processing non-speech or mixed-source audio. In this paper, we first introduce MixFake, a large-scale benchmark dataset designed to simulate diverse acoustic environments with varying SNR levels and mixed authenticity components. To address the "semantic-centric" limitation, we propose a Multi-stream Prompt Tuning framework that injects signal-level priors into SSL backbones. By integrating base, frequency, and texture streams through deep prompt injection, our model effectively captures acoustic artifacts. Experimental results demonstrate that our method significantly outperforms existing baselines, achieving a 0.95% EER in foreground detection and a substantial 7.72% absolute improvement in complex background detection tasks. Our dataset and code are available at https://github.com/saltfish233/MixFake.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。