首个完全开源的韩英双语大模型,用合成数据训练出媲美主流模型的表现。
KORMo: Korean Open Reasoning Model for Everyone
- 从零训练108亿参数韩英双语模型,68.74%韩语数据为合成数据
- 合成数据经精心设计后未引发训练崩溃,性能接近主流多语言模型
- 适合低资源语言研究者复现,推动开放可溯的多语言模型发展
本文首次系统性探索构建完全开源的非英语大型语言模型,聚焦韩语,基于以合成数据为主的韩英语料库进行训练。我们提出KORMo-10B,一个10.8B参数模型,在韩语占比达68.74%的合成数据上从头训练。通过系统实验表明,经过平衡语言覆盖与多样指令风格筛选的合成数据不会导致大规模预训练中的不稳定或性能退化。该模型在推理、知识和指令遵循等多类基准测试中表现与当前主流开源多语言基线相当。实验揭示两个关键发现:(1) 合成数据可在长周期预训练中稳定支持模型不崩溃;(2) 双语指令微调可实现接近母语水平的韩语推理与话语连贯性。本工作全面公开数据、代码、训练配方与日志,建立面向低资源场景的合成数据驱动全开源模型(FOM)透明框架,为未来多语言大模型研究提供可复现范例。
原文摘要 · Abstract (English)
This work presents the first large-scale investigation into constructing a fully open bilingual large language model (LLM) for a non-English language, specifically Korean, trained predominantly on synthetic data. We introduce KORMo-10B, a 10.8B-parameter model trained from scratch on a Korean-English corpus in which 68.74% of the Korean portion is synthetic. Through systematic experimentation, we demonstrate that synthetic data, when carefully curated with balanced linguistic coverage and diverse instruction styles, does not cause instability or degradation during large-scale pretraining. Furthermore, the model achieves performance comparable to that of contemporary open-weight multilingual baselines across a wide range of reasoning, knowledge, and instruction-following benchmarks. Our experiments reveal two key findings: (1) synthetic data can reliably sustain long-horizon pretraining without model collapse, and (2) bilingual instruction tuning enables near-native reasoning and discourse coherence in Korean. By fully releasing all components including data, code, training recipes, and logs, this work establishes a transparent framework for developing synthetic data-driven fully open models (FOMs) in low-resource settings and sets a reproducible precedent for future multilingual LLM research.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。