让Qwen3模型用韩语进行原生推理,提升韩语逻辑能力。
Making Qwen3 Think in Korean with Reinforcement Learning
- 分两阶段微调:先用韩语推理数据集监督微调,再用自研强化学习算法优化
- 在数学和编程等高级推理任务上显著提升,韩语思维链全程保持一致
- 引入裁判模型校准奖励信号,解决强化学习中的奖励欺骗与策略崩溃问题
我们提出一种两阶段微调方法,使大语言模型 Qwen3-14B 实现原生韩语推理。第一阶段在高质量韩语推理数据集上进行监督微调(SFT),显著提升韩语逻辑推理能力,并带来通用推理性能的改善。第二阶段采用定制化的组相对策略优化(GRPO)强化学习算法,进一步增强韩语推理对齐与整体解题能力。针对GRPO训练中常见的奖励欺骗与策略崩溃问题,引入判别器模型校准奖励信号,实现稳定学习并获得持续性能提升。最终模型在数学和编程等高级推理基准测试中表现优异,同时保持知识与语言水平,内部思维链完全以韩语生成。
原文摘要 · Abstract (English)
We present a two-stage fine-tuning approach to make the large language model Qwen3 14B "think" natively in Korean. In the first stage, supervised fine-tuning (SFT) on a high-quality Korean reasoning dataset establishes a strong foundation in Korean logical reasoning, yielding notable improvements in Korean-language tasks and even some gains in general reasoning ability. In the second stage, we employ reinforcement learning with a customized Group Relative Policy Optimization (GRPO) algorithm to further enhance both Korean reasoning alignment and overall problem-solving performance. We address critical stability challenges in GRPO training - such as reward hacking and policy collapse - by introducing an oracle judge model that calibrates the reward signal. Our approach achieves stable learning (avoiding the collapse observed in naive GRPO) and leads to steady, incremental performance gains. The final RL-tuned model demonstrates substantially improved results on advanced reasoning benchmarks (particularly math and coding tasks) while maintaining knowledge and language proficiency, successfully conducting its internal chain-of-thought entirely in Korean.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。