arXiv:2604.01330cs.SDcs.AI2026-04中稿 · WCCI CEC 2026被引 5

用进化算法融合多个深度伪造语音检测器,兼顾准确率与系统轻量化。

Evolutionary Multi-Objective Fusion of Deepfake Speech Detectors

  • 基于NSGA-II算法,同时优化检测误差和模型复杂度
  • 在ASVspoof 5数据集上实现2.37% EER,参数量减半
  • 提供多种精度-效率权衡方案,适合实际部署场景

尽管基于大尺度自监督学习(SSL)模型的深度伪造语音检测器已具备高精度,但采用传统集成融合方法进一步提升鲁棒性时,常导致系统过于庞大且收益递减。为此,本文提出一种基于进化多目标优化的分数融合框架,联合最小化检测错误率与系统复杂度。我们采用NSGA-II优化两种编码方式:二进制编码用于选择检测器进行平均,实数编码则优化各检测器权重以实现加权求和。在包含36个SSL检测器的ASVspoof 5数据集上的实验表明,所获帕累托前沿优于简单平均和逻辑回归基线。实数编码变体达到2.37% EER(0.0684 minDCF),性能媲美当前最优方案,同时参数量仅为一半。该方法还生成多样化的权衡解,支持按需选择精度与计算成本的平衡方案。

原文摘要 · Abstract (English)

While deepfake speech detectors built on large self-supervised learning (SSL) models achieve high accuracy, employing standard ensemble fusion to further enhance robustness often results in oversized systems with diminishing returns. To address this, we propose an evolutionary multi-objective score fusion framework that jointly minimizes detection error and system complexity. We explore two encodings optimized by NSGA-II: binary-coded detector selection for score averaging and a real-valued scheme that optimizes detector weights for a weighted sum. Experiments on the ASVspoof 5 dataset with 36 SSL-based detectors show that the obtained Pareto fronts outperform simple averaging and logistic regression baselines. The real-valued variant achieves 2.37% EER (0.0684 minDCF) and identifies configurations that match state-of-the-art performance while significantly reducing system complexity, requiring only half the parameters. Our method also provides a diverse set of trade-off solutions, enabling deployment choices that balance accuracy and computational cost.

深度伪造检测多目标优化模型融合轻量化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。