arXiv:2510.04230cs.CL2025-10被引 10

用中英混合推理提升多语言模型能力,韩语效果显著

Pushing on Multilingual Reasoning Models with Language-Mixed Chain-of-Thought

  • 设计中英混合思维链,以英文为锚点减少翻译误差
  • 韩语数据集含579万条提示,370万条长推理,35B模型达64.0分新高
  • 小模型也提升18.6分,适合多语言推理研究者使用

近期前沿模型采用长思维链在上下文中探索解空间,实现更强性能。尽管许多工作致力于知识蒸馏以构建更小但高效的模型,但多数集中在英语,对语言特异性推理了解甚少。为此,我们提出**语言混合思维链**(Language-Mixed CoT),通过在英语与目标语言间切换,以英语为锚点强化推理,同时最小化翻译偏差。以韩语为例,我们构建了**Yi-Sang**数据集:包含579万条来自网络问答、考试、理工科和代码的韩语原生提示,370万条由Qwen3-32B生成的长推理轨迹,并筛选出26万条高价值子集。我们在六个模型家族(Qwen2.5、Llama-3.1、Gemma-3等)中训练了九个规模从4B到35B的模型。最佳模型KO-REAson-35B达到64.0±2.5的平均分,9项基准中5项排名第一,其余第二。小模型和中型模型也获得显著提升,九项评测平均提高18.6分。消融实验表明,语言混合思维链优于单语思维链,并带来跨语言与多模态性能增益。我们开源了数据构建流程、评估系统、数据集与模型,推动语言特异性推理研究。数据与模型下载:https://huggingface.co/KOREAson。

原文摘要 · Abstract (English)

Recent frontier models employ long chain-of-thought reasoning to explore solution spaces in context and achieve stonger performance. While many works study distillation to build smaller yet capable models, most focus on English and little is known about language-specific reasoning. To bridge this gap, we first introduct **Language-Mixed CoT**, a reasoning schema that switches between English and a target language, using English as an anchor to excel in reasoning while minimizing translation artificats. As a Korean case study, we curate **Yi-Sang**: 5.79M native-Korean prompts from web Q&A, exams, STEM, and code; 3.7M long reasoning traces generated from Qwen3-32B; and a targeted 260k high-yield subset. We train ninve models (4B-35B) across six families (Qwen2.5, Llama-3.1, Gemma-3, etc). Our best model, **KO-REAson-35B**, achieves state-of-the-art performance, with the highest overall average score (64.0 \pm 25), ranking first on 5/9 benchmarks and second on the remainder. Samller and mid-sized models also benefit substantially, with an average improvement of +18.6 points across teh evaluated nine benchmarks. Ablations show **Language-Mixed CoT** is more effective than monolingual CoT, also resulting in cross-lingual and mult-modal performance gains. We release our data-curation pipeline, evaluation system, datasets, and models to advance research on language-specific reasoning. Data and model collection: https://huggingface.co/KOREAson.

多语言推理思维链韩语模型数据集

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。