大模型答非英文题时偏用英语推理,虽准确率高但易因翻译出错
The Reasoning Lingua Franca: A Double-Edged Sword for Multilingual AI
- 对比模型在母语和英语中的推理过程,发现英语推理更高效
- 复杂任务中英语推理准确率更高,差距随难度增加
- 依赖翻译易引入错误,适合多语言场景的模型需优化
大型推理模型(LRMs)在数学、科学等问答任务中表现优异,但其多语言推理能力尚未充分探索。当面对非英语问题时,LRMs常默认使用英语进行推理,引发可解释性及语言文化细节处理的担忧。我们系统比较了模型在英语与问题母语中的推理表现,涵盖MGSM和GPQA Diamond两个任务。除了评估答案准确率,还分析了推理轨迹中的认知特征。结果发现,英语推理轨迹中这些认知行为显著更多,且整体准确率更高,尤其在复杂任务中差距扩大。然而,这种以英语为中心的策略存在关键缺陷——‘翻译迷失’:翻译步骤可能引入本可避免的错误。
原文摘要 · Abstract (English)
Large Reasoning Models (LRMs) achieve strong performance on mathematical, scientific, and other question-answering tasks, but their multilingual reasoning abilities remain underexplored. When presented with non-English questions, LRMs often default to reasoning in English, raising concerns about interpretability and the handling of linguistic and cultural nuances. We systematically compare an LRM's reasoning in English versus the language of the question. Our evaluation spans two tasks: MGSM and GPQA Diamond. Beyond measuring answer accuracy, we also analyze cognitive attributes in the reasoning traces. We find that English reasoning traces exhibit a substantially higher presence of these cognitive behaviors, and that reasoning in English generally yields higher final-answer accuracy, with the performance gap increasing as tasks become more complex. However, this English-centric strategy is susceptible to a key failure mode - getting "Lost in Translation," where translation steps lead to errors that would have been avoided by reasoning in the language of the question.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。