arXiv:2509.12668eess.AS2025-09被引 3

多阶段融合提升防欺骗语音识别准确率

Investigating the Potential of Multi-Stage Score Fusion in Spoofing-Aware Speaker Verification

  • 分阶段融合语音验证与反欺骗系统输出
  • 在SASV2022数据集上达1.30%错误率
  • 适合关注语音安全与对抗攻击的研究者

尽管自动语音验证(ASV)性能不断提升,其对欺骗攻击的脆弱性仍是重大挑战。本文研究将ASV与反欺骗(CM)子系统整合到模块化防欺骗语音验证(SASV)框架中。不同于传统单阶段得分融合方法,我们探索多阶段策略,分步利用ASV和CM系统输出。基于ECAPA-TDNN(ASV)和AASIST(CM)子系统,采用支持向量机与逻辑回归分类器实现SASV。第二阶段将两者输出与原始得分结合,优化融合后端分类器。此外,引入另一辅助得分(RawGAT,CM)以进一步提升性能。该方法在SASV2022评测数据集上取得1.30%等错误率(EER),相较基线系统相对提升24%。

原文摘要 · Abstract (English)

Despite improvements in automatic speaker verification (ASV), vulnerability against spoofing attacks remains a major concern. In this study, we investigate the integration of ASV and countermeasure (CM) subsystems into a modular spoof-aware speaker verification (SASV) framework. Unlike conventional single-stage score-level fusion methods, we explore the potential of a multi-stage approach that utilizes the ASV and CM systems in multiple stages. By leveraging ECAPA-TDNN (ASV) and AASIST (CM) subsystems, we consider support vector machine and logistic regression classifiers to achieve SASV. In the second stage, we integrate their outputs with the original score to revise fusion back-end classifiers. Additionally, we incorporate another auxiliary score from RawGAT (CM) to further enhance our SASV framework. Our approach yields an equal error rate (EER) of 1.30% on the evaluation dataset of the SASV2022 challenge, representing a 24% relative improvement over the baseline system.

语音验证反欺骗多阶段融合

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。