用语音内容验证提升说话人识别准确率,系统表现优异。
The SVASR System for Text-dependent Speaker Verification (TdSV) AAIC Challenge 2024
- 用快速Conformer模型验证语音内容,过滤错误匹配
- 融合wav2vec-BERT与ReDimNet特征,提升识别精度
- 在2024挑战赛中排名第二,适合高安全场景应用
本文提出一种高效精准的文本依赖说话人验证(TDSV)系统,以满足高性能生物特征识别需求。系统采用基于Fast-Conformer的ASR模块验证语音内容,有效过滤目标错误(TW)和冒名错误(IW)测试样本。在说话人验证阶段,通过融合wav2vec-BERT与ReDimNet模型提取的说话人嵌入,构建统一的说话人表征。该系统在TDSV 2024挑战赛测试集上取得0.0452的归一化最小代价函数(min-DCF),排名第二,展现出出色的准确性与鲁棒性平衡能力。
原文摘要 · Abstract (English)
This paper introduces an efficient and accurate pipeline for text-dependent speaker verification (TDSV), designed to address the need for high-performance biometric systems. The proposed system incorporates a Fast-Conformer-based ASR module to validate speech content, filtering out Target-Wrong (TW) and Impostor-Wrong (IW) trials. For speaker verification, we propose a feature fusion approach that combines speaker embeddings extracted from wav2vec-BERT and ReDimNet models to create a unified speaker representation. This system achieves competitive results on the TDSV 2024 Challenge test set, with a normalized min-DCF of 0.0452 (rank 2), highlighting its effectiveness in balancing accuracy and robustness.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。