arXiv:2605.01905cs.SDcs.CL2026-05中稿 · Odyssey 2026

用预训练模型和边界损失提升语音语言识别准确率

Spoken Language Identification with Pre-trained Models and Margin Loss

  • 基于ECAPA-TDNN提取特征,结合边界损失增强语言表征区分度
  • 在Tidy-X数据集上达到85.95%宏准确率与90.96%微准确率
  • 显著降低说话人干扰,适合语音识别与验证场景应用

针对TidyLang Challenge 2026提出的说话人控制语音语言识别任务,本文提出一种基于预训练模型与边界损失的语言识别方法。该方法采用预训练的ECAPA-TDNN作为特征编码器,并引入边界损失以增强语言表征的判别能力,从而提高类间可分性并减少说话人等非语言因素的干扰。在Tidy-X数据集上的实验结果表明,该方法在语言识别任务中取得85.95%的宏准确率和90.96%的微准确率,在验证任务中实现17.08%的等错误率(EER)。相比官方基线,宏准确率提升45.7%,微准确率提升15.2%,EER降低约50.8%,充分验证了方法的有效性。代码将公开于https://github.com/PunkMale/TidyLang2026。

原文摘要 · Abstract (English)

For the speaker-controlled spoken language identification task proposed in the TidyLang Challenge 2026, this paper proposes a language identification method based on pre-trained models and margin-based losses. The proposed method adopts a pre-trained ECAPA-TDNN as the feature encoder and incorporates margin-based losses to enhance the discriminative ability of language representations, thereby improving inter-class separability and reducing the interference of non-linguistic factors such as speaker characteristics. Experimental results on the Tidy-X dataset show that the proposed method achieves 85.95% macro accuracy and 90.96% micro accuracy on the language identification task and 17.08% equal error rate (EER) on the verification task. Compared with the official baseline, the macro accuracy improves by 45.7%, the micro accuracy improves by 15.2%, and the EER is reduced by approximately 50.8%, demonstrating the effectiveness of the proposed method. The code will be released at https://github.com/PunkMale/TidyLang2026.

语言识别预训练模型边界损失语音处理

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。