arXiv:2602.01722eess.AS2026-02

通过非线性融合提升语音伪造检测与声纹识别联合性能

Joint Optimization of ASV and CM tasks: BTUEF Team's Submission for WildSpoof Challenge

  • 用非线性融合整合现成声纹与反伪造模型,显式建模交互关系
  • 在野声场景下达到0.0515的a-DCF最优表现(进度评估集)
  • 适合关注语音安全、对抗攻击防御的研究者和工程师

针对对抗性攻击下的鲁棒性挑战,本文研究了一种模块化联合处理声纹验证(ASV)与语音伪造检测(CM)的框架。该框架通过非线性融合,有效复用公开的ASV与CM系统,显式建模两者交互,并采用依赖运行条件的可训练a-DCF损失进行优化。实验采用ECAPA-TDNN与ReDimNet作为ASV嵌入提取器,SSL-AASIST作为CM模型,在有无对WildSpoof SASV训练数据微调的条件下进行了测试。结果表明,基于ReDimNet的ASV嵌入与微调后的SSL-AASIST表示组合效果最佳,在进度评估集上a-DCF为0.0515,在最终评估集上为0.2163。

原文摘要 · Abstract (English)

Spoofing-aware speaker verification (SASV) jointly addresses automatic speaker verification and spoofing countermeasures to improve robustness against adversarial attacks. In this paper, we investigate our recently proposed modular SASV framework that enables effective reuse of publicly available ASV and CM systems through non-linear fusion, explicitly modeling their interaction, and optimization with an operating-condition-dependent trainable a-DCF loss. The framework is evaluated using ECAPA-TDNN and ReDimNet as ASV embedding extractors and SSL-AASIST as the CM model, with experiments conducted both with and without fine-tuning on the WildSpoof SASV training data. Results show that the best performance is achieved by combining ReDimNet-based ASV embeddings with fine-tuned SSL-AASIST representations, yielding an a-DCF of 0.0515 on the progress evaluation set and 0.2163 on the final evaluation set.

声纹验证语音伪造对抗攻击联合优化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。