arXiv:2605.14896cs.SDcs.LG2026-05

基于轻量模型与集成学习,实现高精度语音身份验证。

Text-Dependent Speaker Verification (TdSV) Challenge 2024: Team Naive System Report

论文配图:Text-Dependent Speaker Verification (TdSV) Challenge 2024: Team Naive System Report
图 1 · 摘自论文原文
  • 复用ResNet-TDNN等预训练模型,结合自研EfficientNet-A0增强适应性。
  • 在挑战赛中达到MinDCF 0.0461、EER 1.3%的优异性能。
  • 适合资源有限但需快速部署的语音验证场景。

本文介绍2024年文本依赖语音验证(TdSV)挑战赛的参赛系统。该系统在测试中取得最小检测成本函数(MinDCF)0.0461、等错误率(EER)1.3%的成绩。方法上,采用在VoxCeleb数据集上预训练的先进神经网络架构,如ResNet-TDNN和NeXt-TDNN,并针对挑战赛数据集设计了轻量级高效模型EfficientNet-A0以提升适配效果。系统融合多种先进模型结构、大规模数据增强及优化超参数配置,显著提升了文本依赖语音验证性能。结果表明,多模型集成学习在说话人与语句验证任务中均具有效力。

原文摘要 · Abstract (English)

This paper presents a system for the 2024 Text-Dependent Speaker Verification (TdSV) Challenge. The system achieved a Minimum Detection Cost Function (MinDCF) of 0.0461 and an Equal Error Rate (EER) of 1.3\%. Our approach focused on adapting existing state-of-the-art neural networks, ResNet-TDNN and NeXt-TDNN, originally trained on the VoxCeleb dataset. This strategy was chosen because of the limited challenge duration and the available resources at the time. In addition, we designed a lightweight and resource-efficient model, EfficientNet-A0, trained specifically on the challenge dataset to improve adaptation and strengthen the ensemble approach. Our system combines advanced neural architectures, extensive data augmentation, and optimised hyperparameters. These components helped achieve strong performance in text-dependent speaker verification. The results also demonstrate the effectiveness of multi-model ensemble learning for both speaker and phrase verification.

语音验证深度学习模型集成轻量化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。