TalTech团队在2025年国际语音挑战赛中夺冠,构建了高效多语言语音识别与语言识别系统。
TalTech Systems for the Interspeech 2025 ML-SUPERB 2.0 Challenge
- 采用共享编码器的混合语言识别架构,结合预训练语言嵌入与轻量级语言模型。
- 针对不同语言数据量,选用Fine-tuned SeamlessM4T、MMS-1B适配器或零样本模型,实现最优性能。
- 系统在挑战赛中综合得分排名第一,适合多语言语音处理研究者参考。
本文介绍塔林理工大学为参加2025年Interspeech ML-SUPERB 2.0挑战赛所开发的语言识别与多语言语音识别系统。系统采用混合语言识别架构,包含一个预训练语言嵌入模型和一个具有跨语言共享编码器的轻量级语音识别模型,搭配语言特定的二元语法语言模型。语音识别部分使用三个模型:根据训练数据可用性和保留数据上的表现,每种语言仅应用一个模型。模型组合包括SeamlessM4T的微调版本、带有自定义语言适配器的MMS-1B-all,以及MMS-zeroshot。该系统在挑战赛中获得总体最高分。
原文摘要 · Abstract (English)
This paper describes the language identification and multilingual speech recognition system developed at Tallinn University of Technology for the Interspeech 2025 ML-SUPERB 2.0 Challenge. A hybrid language identification system is used, consisting of a pretrained language embedding model and a light-weight speech recognition model with a shared encoder across languages and language-specific bigram language models. For speech recognition, three models are used, where only a single model is applied for each language, depending on the training data availability and performance on held-out data. The model set consists of a finetuned version of SeamlessM4T, MMS-1B-all with custom language adapters and MMS-zeroshot. The system obtained the top overall score in the challenge.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。