USTC-KXDIGIT系统在语音伪造检测中表现优异,兼顾闭包与开包场景的鲁棒性。
USTC-KXDIGIT System Description for ASVspoof5 Challenge
- 结合手工特征与自监督模型生成嵌入,提升特征表达能力
- 通过数据增强和语音转换扩充伪造样本,提高模型泛化性
- 融合多模型输出得分,实现更精准的语音伪造检测与抗欺骗验证
本文描述了提交至ASVspoof5挑战赛第1赛道(语音深度伪造检测)和第2赛道(抗欺骗自动说话人验证,SASV)的USTC-KXDIGIT系统。第1赛道涵盖多种处理算法及闭包与开包条件。系统由前端特征提取器与后端分类器级联构成,重点在于嵌入工程优化与后端分类器泛化能力提升。嵌入工程分别基于手工特征(闭包)与自监督模型语音表征(开包)。为应对多样化对抗条件,采用增强训练集并使用语音转换技术合成假音频以丰富生成算法。通过激活集成与多系统分数融合,获得最终决策得分。评估阶段,闭包条件下minDCF为0.3948,EER为14.33%;开包条件下minDCF为0.0750,EER为2.59%,展现出良好鲁棒性。第2赛道延续第1赛道的CM系统,并融合基于CNN的ASV系统,闭包条件min-aDCF达0.2814,开包条件为0.0756,性能优越。
原文摘要 · Abstract (English)
This paper describes the USTC-KXDIGIT system submitted to the ASVspoof5 Challenge for Track 1 (speech deepfake detection) and Track 2 (spoofing-robust automatic speaker verification, SASV). Track 1 showcases a diverse range of technical qualities from potential processing algorithms and includes both open and closed conditions. For these conditions, our system consists of a cascade of a frontend feature extractor and a back-end classifier. We focus on extensive embedding engineering and enhancing the generalization of the back-end classifier model. Specifically, the embedding engineering is based on hand-crafted features and speech representations from a self-supervised model, used for closed and open conditions, respectively. To detect spoof attacks under various adversarial conditions, we trained multiple systems on an augmented training set. Additionally, we used voice conversion technology to synthesize fake audio from genuine audio in the training set to enrich the synthesis algorithms. To leverage the complementary information learned by different model architectures, we employed activation ensemble and fused scores from different systems to obtain the final decision score for spoof detection. During the evaluation phase, the proposed methods achieved 0.3948 minDCF and 14.33% EER in the close condition, and 0.0750 minDCF and 2.59% EER in the open condition, demonstrating the robustness of our submitted systems under adversarial conditions. In Track 2, we continued using the CM system from Track 1 and fused it with a CNN-based ASV system. This approach achieved 0.2814 min-aDCF in the closed condition and 0.0756 min-aDCF in the open condition, showcasing superior performance in the SASV system.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。