arXiv:2409.16077cs.SDcs.AI2024-09被引 24

用专家混合模型提升语音伪造检测的泛化能力

Leveraging Mixture of Experts for Improved Speech Deepfake Detection

  • 采用专家混合架构,动态分配不同专家权重以适应多样输入
  • 在多个数据集上表现优于传统单模型和集成方法
  • 轻量级门控机制适合持续更新,适合应对新型伪造技术

语音深度伪造对个人安全和内容真实性构成重大威胁。现有检测系统面临的一个主要挑战是泛化能力不足,难以识别跨数据集的虚假信号。本文提出一种基于专家混合(Mixture of Experts)架构的新方法,以提升语音深伪检测性能。该框架因能针对不同输入类型进行专业化处理并高效应对数据变异性,特别适用于语音深伪检测任务。相比传统单模型或集成方法,该方法具备更优的泛化能力和对未见数据的适应性。其模块化结构支持可扩展更新,有助于应对不断演进的伪造技术,同时保持高检测准确率。我们设计了一种高效轻量的门控机制,动态为每个输入分配专家权重,从而优化检测性能。多数据集实验验证了所提方法的有效性与潜力。

原文摘要 · Abstract (English)

Speech deepfakes pose a significant threat to personal security and content authenticity. Several detectors have been proposed in the literature, and one of the primary challenges these systems have to face is the generalization over unseen data to identify fake signals across a wide range of datasets. In this paper, we introduce a novel approach for enhancing speech deepfake detection performance using a Mixture of Experts architecture. The Mixture of Experts framework is well-suited for the speech deepfake detection task due to its ability to specialize in different input types and handle data variability efficiently. This approach offers superior generalization and adaptability to unseen data compared to traditional single models or ensemble methods. Additionally, its modular structure supports scalable updates, making it more flexible in managing the evolving complexity of deepfake techniques while maintaining high detection accuracy. We propose an efficient, lightweight gating mechanism to dynamically assign expert weights for each input, optimizing detection performance. Experimental results across multiple datasets demonstrate the effectiveness and potential of our proposed approach.

语音伪造专家混合检测

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。