arXiv:2509.17585cs.SD2025-09中稿 · @ IEEE WIFS 2025被引 3

用注意力门控融合多个专家模型,提升语音伪造检测鲁棒性

Attention-based Mixture of Experts for Robust Speech Deepfake Detection

  • 通过注意力机制动态加权多个专家模型的输出
  • 在SAFE挑战中所有任务均排名第一,超越现有方法
  • 适合需要高可靠语音安全检测的场景

AI生成语音正广泛应用于虚拟助手、无障碍工具等场景,但亦被用于身份冒用、虚假信息和生物特征欺骗等恶意行为。随着语音深伪技术逼近真实人类语音,亟需可靠的检测方法。本文提出ISPL在IH&MMSec 2025 SAFE挑战中的参赛方案,系统在所有任务中排名第一。该方法基于专家混合(Mixture of Experts)架构,融合多个前沿检测器,通过注意力门控网络根据输入语音信号动态调整各专家权重。每个专家通过归纳偏置学习输入的不同互补特征,实现对训练数据的专门化理解。实验表明,该方法在多个数据集上优于现有技术,并在SAFE挑战中验证了其卓越性能。

原文摘要 · Abstract (English)

AI-generated speech is becoming increasingly used in everyday life, powering virtual assistants, accessibility tools, and other applications. However, it is also being exploited for malicious purposes such as impersonation, misinformation, and biometric spoofing. As speech deepfakes become nearly indistinguishable from real human speech, the need for robust detection methods and effective countermeasures has become critically urgent. In this paper, we present the ISPL's submission to the SAFE challenge at IH&MMSec 2025, where our system ranked first across all tasks. Our solution introduces a novel approach to audio deepfake detection based on a Mixture of Experts architecture. The proposed system leverages multiple state-of-the-art detectors, combining their outputs through an attention-based gating network that dynamically weights each expert based on the input speech signal. In this design, each expert develops a specialized understanding of the shared training data by learning to capture different complementary aspects of the same input through inductive biases. Experimental results indicate that our method outperforms existing approaches across multiple datasets. We further evaluate and analyze the performance of our system in the SAFE challenge.

语音伪造专家混合注意力机制

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。