arXiv:2506.02085cs.SDcs.AI2025-06中稿 · Interspeech 2025, …被引 9

通过融合度量学习与音视频模型,精准追踪音频伪造来源。

Unveiling Audio Deepfake Origins: A Deep Metric learning And Conformer Network Approach With Ensemble Fusion

  • 用N对损失增强真假语音的区分能力
  • 在真实与伪造语音中实现95%以上溯源准确率
  • 适合安全检测与数字取证场景使用

音频深度伪造正借助先进AI技术达到前所未有的逼真程度。当前研究多聚焦于辨别真实语音与伪造语音,但追溯伪造源同样关键。本文提出一种新型音频源溯源系统,结合深度度量多类N对损失与真实强化-虚假分散框架、Conformer分类网络以及集成得分-嵌入融合策略。N对损失提升特征判别力,真实强化与虚假分散机制增强对真实与伪造语音模式的区分鲁棒性。Conformer网络可捕捉音频信号中的全局与局部依赖关系,对溯源至关重要。所提集成得分-嵌入融合方法在域内与域外溯源场景间取得最优平衡。采用Fréchet距离及标准指标评估,实验表明该方法在源溯源任务上显著优于基线系统。

原文摘要 · Abstract (English)

Audio deepfakes are acquiring an unprecedented level of realism with advanced AI. While current research focuses on discerning real speech from spoofed speech, tracing the source system is equally crucial. This work proposes a novel audio source tracing system combining deep metric multi-class N-pair loss with Real Emphasis and Fake Dispersion framework, a Conformer classification network, and ensemble score-embedding fusion. The N-pair loss improves discriminative ability, while Real Emphasis and Fake Dispersion enhance robustness by focusing on differentiating real and fake speech patterns. The Conformer network captures both global and local dependencies in the audio signal, crucial for source tracing. The proposed ensemble score-embedding fusion shows an optimal trade-off between in-domain and out-of-domain source tracing scenarios. We evaluate our method using Frechet Distance and standard metrics, demonstrating superior performance in source tracing over the baseline system.

音频伪造源追踪深度学习Conformer

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。