研究推理模型的语言混用现象及其影响,发现强制使用特定语言可提升准确率。
Language Mixing in Reasoning Language Models: Patterns, Impact, and Internal Causes
- 系统分析15种语言、7类难度下的推理语言混用模式。
- 强制用拉丁或汉字推理可显著提高模型准确率。
- 模型内部表征与推理文本的书写系统高度一致。
推理语言模型(RLMs)通过链式思维生成结构化中间步骤,在复杂任务中表现优异。然而,其输出中常出现与提示语不同语言的词汇(即语言混用),影响性能,但具体影响尚存争议。本文首次对RLMs中的语言混用进行系统研究,覆盖15种语言、7个任务难度等级和18个学科领域,揭示三者对语言混用的共同影响。进一步发现,推理语言的选择显著影响性能:通过约束解码强制模型使用拉丁语或汉字进行推理,能明显提升准确率。此外,推理过程的书写系统与模型内部表示的书写系统高度匹配,表明语言混用反映了模型潜在的处理偏好。研究结果为优化多语言推理提供了可操作洞见,并为控制推理语言以构建更可解释、更灵活的推理模型开辟新方向。
原文摘要 · Abstract (English)
Reasoning language models (RLMs) excel at complex tasks by leveraging a chain-of-thought process to generate structured intermediate steps. However, language mixing, i.e., reasoning steps containing tokens from languages other than the prompt, has been observed in their outputs and shown to affect performance, though its impact remains debated. We present the first systematic study of language mixing in RLMs, examining its patterns, impact, and internal causes across 15 languages, 7 task difficulty levels, and 18 subject areas, and show how all three factors influence language mixing. Moreover, we demonstrate that the choice of reasoning language significantly affects performance: forcing models to reason in Latin or Han scripts via constrained decoding notably improves accuracy. Finally, we show that the script composition of reasoning traces closely aligns with that of the model's internal representations, indicating that language mixing reflects latent processing preferences in RLMs. Our findings provide actionable insights for optimizing multilingual reasoning and open new directions for controlling reasoning languages to build more interpretable and adaptable RLMs.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。