arXiv:2505.18273eess.AS2025-05被引 3

用评分门控融合防伪分数与声纹特征,提升语音验证鲁棒性

ATMM-SAGA: Alternating Training for Multi-Module with Score-Aware Gated Attention SASV system

  • 通过门控机制动态融合防伪评分与声纹嵌入
  • 在ASVspoof2019数据集上达到2.18%的EER和0.0480的a-DCF
  • 适合需要抵御语音伪造攻击的高安全场景

自动说话人验证(ASV)系统旨在判断给定语音是否属于注册说话人。为确保系统可靠性,本文提出一种抗欺骗的说话人验证(SASV)系统,采用评分感知门控注意力(SAGA)融合方案,整合预训练防伪模块(CM)的得分与预训练声纹验证(ASV)的说话人嵌入。实验基于ASVspoof2019逻辑访问数据集,结果表明,该SASV系统在开发集上实现2.31%的SASV等错误率(SASV-EER)和0.0603的通用检测代价函数(a-DCF),在测试集上分别达到2.18%和0.0480。所用模型包括AASIST和ECAPA-TDNN。

原文摘要 · Abstract (English)

The objective of automatic speaker verification (ASV) systems is to determine whether a given test speech utterance corresponds to a claimed enrolled speaker. These systems have a wide range of applications, and ensuring their reliability is crucial. In this paper, we propose a spoofing-robust automatic speaker verification (SASV) system employing a score-aware gated attention (SAGA) fusion scheme, integrating scores from a pre-trained countermeasure (CM) with speaker embeddings from a pre-trained ASV. Specifically, we employ the AASIST and ECAPA-TDNN models. SAGA acts as an adaptive gating mechanism, where the CM score determines how strongly ASV embeddings influence the final SASV decision. Experiments on the ASVspoof2019 logical access dataset demonstrate that the proposed SASV system achieves an SASV equal error rate (SASV-EER) and agnostic detection cost function (a-DCF) of 2.31%, 0.0603 for the development set and 2.18%, 0.0480 for the evaluation set.

说话人验证防伪攻击门控机制音视频安全

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。