让大模型用日语推理可行但效果有限,且文化任务表现更差。
Cost of Reasoning in non-English Languages: A Case Study on Japanese

- 用GRPO持续微调日语版Qwen-3,实现日语推理控制。
- 日语推理模型在编程、数学等任务上与英语强基线持平。
- 日语推理在文化相关任务上反而不如基线,需额外优化。
推理语言模型(RLMs)在英语环境下表现最佳,因其训练数据最丰富。然而,推理过程对模型可解释性和安全性至关重要,对用户和开发者均有实用价值。因此,开发能在用户选择语言中进行推理的模型极具意义。本文研究了在日语中进行推理的可行性,构建了基于Qwen-3-8B持续预训练的日语推理模型Qwen-3-Swallow-8B,采用GRPO方法训练,并在编码、数学和科学基准上评估其表现。结果表明,通过持续预训练与GRPO可实现日语推理控制,但性能最多仅达到强英语推理基线水平。进一步在日语文化基准上测试发现,该模型表现劣于基线,说明日语推理不会自动提升文化相关任务性能。
原文摘要 · Abstract (English)
Reasoning Language Models (RLMs) achieve their strongest performance when they reason in English, the language for which reasoning-oriented training data is most abundant. However, reasoning trace is a clue for model interpretability and safety, and useful in practice for both the model users and for model developers. Thus, it is desirable to be able to develop a model that reasons in a language of the user's choice, while still maintaining strong reasoning performance. To this end, we study the feasibility of training a model that reasons in Japanese. We develop a Japanese-reasoning variant of Qwen-3-Swallow-8B, which is a Japanese LLM continually pretrained from Qwen-3-8B, with GRPO and evaluate it across coding, math, and science benchmarks. The study shows that reasoning-language control is feasible by training a Japanese continually pretrained model with GRPO. However, its performance is at best on par with strong English-reasoning baselines on several benchmarks. We also evaluate the trained model on Japanese cultural benchmarks and observe that the model's performance is worse than the baseline models, suggesting that the reasoning in Japanese does not immediately improve performance on culturally relevant tasks for free.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。