arXiv:2507.19308cs.SDcs.CL2025-07被引 4

探索对比学习与对话上下文对多语言语音识别鲁棒性的提升作用

The Eloquence team submission for task 1 of MLC-SLM challenge

论文配图:The Eloquence team submission for task 1 of MLC-SLM challenge
图 1 · 摘自论文原文
  • 采用线性投影器与qformer改进多语言语音模型架构
  • 对比学习与扩展对话上下文显著提升识别准确率
  • 适合关注多语言语音系统鲁棒性优化的研究者

本文针对多语言对话语音语言模型(MLC-SLM)挑战赛任务1,研究并实验了三种多语言语音识别方法。首先,通过在不同基础模型上训练线性与qformer投影器,评估官方基线的优劣;其次,基于SLAM-ASR框架训练自定义的多语言线性投影器;最后,探究对比学习与扩展对话上下文对识别鲁棒性的增强效果。实验表明,引入对比学习和更长对话上下文能有效提升多语言场景下的语音识别性能。

原文摘要 · Abstract (English)

In this paper, we present our studies and experiments carried out for the task 1 of the Challenge and Workshop on Multilingual Conversational Speech Language Model (MLC-SLM), which focuses on advancing multilingual conversational speech recognition through the development of speech language models architectures. Given the increasing relevance of real-world conversational data for building robust Spoken Dialogue Systems, we explore three approaches to multilingual ASR. First, we conduct an evaluation of the official baseline to better understand its strengths and limitations, by training two projectors (linear and qformer) with different foundation models. Second we leverage the SLAM-ASR framework to train a custom multilingual linear projector. Finally we investigate the role of contrastive learning and the extended conversational context in enhancing the robustness of recognition.

多语言语音对比学习对话上下文

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。