大模型推理时更倾向用英语,影响多语言任务表现。
Language Matters: How Do Multilingual Input and Reasoning Paths Affect Large Reasoning Models?
- 模型在多语言输入下默认用高资源语言(如英语)推理
- 低资源语言输入时性能下降,但用英语推理能保持表现
- 推理语言影响不同任务:推理类任务降分,文化类任务提分
大型推理模型(LRMs)在多种推理任务中表现优异,但其在多语言环境下的内部推理过程仍不明确。我们提出核心问题:当问题以不同语言呈现时,这些模型实际用哪种语言进行推理?研究发现,尽管经过多语言训练,模型在测试时仍倾向于默认使用高资源语言(如英语)进行推理,无论输入语言为何。若强制模型使用与输入相同的语言推理,性能显著下降,尤其在低资源语言上。相反,使用高资源语言推理则通常能保持性能。我们在多个推理密集型任务(MMMLU、MATH-500)和非推理基准(CulturalBench、LMSYS-toxic)上进行了广泛评估,发现语言选择的影响因任务类型而异:输入语言推理会降低推理类任务表现,但提升文化相关任务表现;安全评估则呈现语言特异性行为。本工作揭示了大模型中的语言偏见,为开发服务多元语言背景用户的更公平模型提供了关键步骤。
原文摘要 · Abstract (English)
Large reasoning models (LRMs) have demonstrated impressive performance across a range of reasoning tasks, yet little is known about their internal reasoning processes in multilingual settings. We begin with a critical question: {\it In which language do these models reason when solving problems presented in different languages?} Our findings reveal that, despite multilingual training, LRMs tend to default to reasoning in high-resource languages (e.g., English) at test time, regardless of the input language. When constrained to reason in the same language as the input, model performance declines, especially for low-resource languages. In contrast, reasoning in high-resource languages generally preserves performance. We conduct extensive evaluations across reasoning-intensive tasks (MMMLU, MATH-500) and non-reasoning benchmarks (CulturalBench, LMSYS-toxic), showing that the effect of language choice varies by task type: input-language reasoning degrades performance on reasoning tasks but benefits cultural tasks, while safety evaluations exhibit language-specific behavior. By exposing these linguistic biases in LRMs, our work highlights a critical step toward developing more equitable models that serve users across diverse linguistic backgrounds.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。