arXiv:2507.18051cs.SDeess.AS2025-07被引 4

多语言对话语音识别与说话人分离系统,显著降低错误率并获挑战赛冠亚军。

The TEA-ASLP System for Multilingual Conversational Speech Recognition and Speech Diarization in MLC-SLM 2025 Challenge

  • 融合语言识别与多语言稀疏专家模型,用CTC输出作为提示提升生成质量。
  • 任务一WER降至9.60%,比基线降低30.8%;任务二时间约束下WER为17.49%。
  • 适用于多语言对话场景,适合关注语音识别与分离技术的开发者和研究者。

本文介绍TEA-ASLP团队提交至MLC-SLM 2025挑战赛的系统,涵盖任务一的多语言对话自动语音识别(ASR)和任务二的说话人分离语音识别。在任务一中,通过集成已知语言识别信息、多语言稀疏专家(MOE)LoRA结构,并利用CTC预测的词元作为提示,增强Ideal-LLM模型,训练数据约180k小时。任务二中,将基线英中文说话人分离模型替换为更合适的纯英语版本。该方法相较基线语音语言模型在任务一中实现30.8%的词错误率(WER)下降,最终获得9.60%的WER;任务二在时间约束下的最小排列WER为17.49%,分别获得两个任务的第一名和第二名。

原文摘要 · Abstract (English)

This paper presents the TEA-ASLP's system submitted to the MLC-SLM 2025 Challenge, addressing multilingual conversational automatic speech recognition (ASR) in Task I and speech diarization ASR in Task II. For Task I, we enhance Ideal-LLM model by integrating known language identification and a multilingual MOE LoRA structure, along with using CTC-predicted tokens as prompts to improve autoregressive generation. The model is trained on approximately 180k hours of multilingual ASR data. In Task II, we replace the baseline English-Chinese speaker diarization model with a more suitable English-only version. Our approach achieves a 30.8% reduction in word error rate (WER) compared to the baseline speech language model, resulting in a final WER of 9.60% in Task I and a time-constrained minimum-permutation WER of 17.49% in Task II, earning first and second place in the respective challenge tasks.

语音识别多语言说话人分离ASR

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。