arXiv:2507.15849cs.CLcs.AI2025-07EMNLP被引 12

语言切换能提升双语大模型推理能力,非无意识现象

The Impact of Language Mixing on Bilingual LLM Reasoning

  • 通过强化学习与可验证奖励发现语言切换关键训练阶段
  • 强制单语解码导致MATH500准确率下降5.6个百分点
  • 轻量级探测器可引导语言切换,提升准确率2.92个百分点

熟练的双语使用者常在对话中主动切换语言。类似地,具备双语强推理能力的大型语言模型(LLMs)也会在思维链中交替使用语言。研究发现,禁止DeepSeek-R1的语言切换会降低准确率,暗示语言切换可能有助于推理。本文研究中文-英文双语推理模型中的语言切换行为,识别出强化学习与可验证奖励(RLVR)是导致语言切换的关键训练阶段。实验表明,强制单语解码使MATH500数据集上的准确率下降5.6个百分点;此外,可训练一个轻量级探测器预测语言切换是否有益,用于指导解码后,准确率提升2.92个百分点。结果表明,语言切换并非多语言训练的副产物,而是一种有策略的推理行为。

原文摘要 · Abstract (English)

Proficient multilingual speakers often intentionally switch languages in the middle of a conversation. Similarly, recent reasoning-focused bilingual large language models (LLMs) with strong capabilities in both languages exhibit language mixing-alternating languages within their chain of thought. Discouraging this behavior in DeepSeek-R1 was found to degrade accuracy, suggesting that language mixing may benefit reasoning. In this work, we study language switching in Chinese-English bilingual reasoning models. We identify reinforcement learning with verifiable rewards (RLVR) as the critical training stage that leads to language mixing. We show that language mixing can enhance reasoning: enforcing monolingual decoding reduces accuracy by 5.6 percentage points on MATH500. Additionally, a lightweight probe can be trained to predict whether a potential language switch would benefit or harm reasoning, and when used to guide decoding, increases accuracy by 2.92 percentage points. Our findings suggest that language mixing is not merely a byproduct of multilingual training, but is a strategic reasoning behavior.

双语模型语言切换推理增强

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。