arXiv:2510.07881cs.CL2025-10被引 3

针对中英混用语音交互模型语言对齐不足问题,提出新评测基准与改进方法。

CS3-Bench: Evaluating and Enhancing Speech-to-Speech LLMs for Mandarin-English Code-Switching

  • 构建中英混用语音指令评测集,发现主流模型性能下降超66%。
  • 引入识别链与关键词标注,知识问答准确率从25.14%提升至46.13%。
  • 适合语音交互、多语言大模型研究者参考,尤其关注语言混合场景。

多模态大语言模型的发展推动了语音到语音交互系统进步。尽管单语自然交互已实现,但现有模型在语言对齐方面存在缺陷。我们提出的中英混用语音评测基准(CS3-Bench)显示,7个主流模型在知识密集型问答任务中相对性能下降高达66%,开放对话中也出现不同程度误解。基于表现严重退化的模型,我们提出数据构建与训练策略,采用识别链(CoR)增强理解,关键词标注(KH)引导生成。该方法使知识准确率从25.14%提升至46.13%,开放对话理解率从64.5%提升至86.5%,显著减少次语言发音错误。CS3-Bench 已开源于 https://huggingface.co/datasets/VocalNet/CS3-Bench。

原文摘要 · Abstract (English)

The advancement of multimodal large language models has accelerated the development of speech-to-speech interaction systems. While natural monolingual interaction has been achieved, we find existing models exhibit deficiencies in language alignment. In our proposed Code-Switching Speech-to-Speech Benchmark (CS3-Bench), experiments on 7 mainstream models demonstrate a relative performance drop of up to 66% in knowledge-intensive question answering and varying degrees of misunderstanding in open-ended conversations. Starting from a model with severe performance deterioration, we propose both data constructions and training approaches to improve the language alignment capabilities, specifically employing Chain of Recognition (CoR) to enhance understanding and Keyword Highlighting (KH) to guide generation. Our approach improves the knowledge accuracy from 25.14% to 46.13%, with open-ended understanding rate from 64.5% to 86.5%, and significantly reduces pronunciation errors in the secondary language. CS3-Bench is available at https://huggingface.co/datasets/VocalNet/CS3-Bench.

语音生成多语言代码混用大模型评测

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。