评测大模型对混语文本的理解与生成能力,发现其仍有明显短板。
Can Large Language Models Understand, Reason About, and Generate Code-Switched Text?
- 构建了16种混语配对的高质量评测集CodeMixQA
- 大模型在混语问答中推理准确率不足,生成文本自然度低
- 适合关注多语言模型鲁棒性的研究者参考
代码切换是多语言交流中的普遍现象,但大语言模型(LLM)在混合语言环境下的鲁棒性仍不明确。本文提出一个全面评估框架,分析LLM在理解、推理和生成混语文本方面的能力。我们构建了CodeMixQA,一个包含16种不同语言组合的高质量人工标注基准,覆盖多个地理区域和代码切换模式,并包含原生文字与转写形式。基于此,我们分析了模型在混语问答任务中的推理行为,揭示其处理混合语言输入的方式。进一步系统评估了模型生成的合成混语文本,重点关注自然度与语义保真度,发现当前生成能力存在明显局限。结果表明,混语环境下推理与生成仍面临持续挑战,为构建更鲁棒的多语言大模型提供可操作洞察。数据集与代码已开源。
原文摘要 · Abstract (English)
Code-switching is a pervasive phenomenon in multilingual communication, yet the robustness of large language models (LLMs) in mixed-language settings remains insufficiently understood. In this work, we present a comprehensive evaluation of LLM capabilities in understanding, reasoning over, and generating code-switched text. We introduce CodeMixQA a novel benchmark with high-quality human annotations, comprising 16 diverse parallel code-switched language-pair variants that span multiple geographic regions and code-switching patterns, and include both original scripts and their transliterated forms. Using this benchmark, we analyze the reasoning behavior of LLMs on code-switched question-answering tasks, shedding light on how models process and reason over mixed-language inputs. We further conduct a systematic evaluation of LLM-generated synthetic code-switched text, focusing on both naturalness and semantic fidelity, and uncover key limitations in current generation capabilities. Our findings reveal persistent challenges in both reasoning and generation under code-switching conditions and provide actionable insights for building more robust multilingual LLMs. We release the dataset and code as open source.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。