102B参数模型助力冷门语言,通过合成数据与强化学习提升性能
Solar Open Technical Report
- 合成4.5万亿高质量语料,解决冷门语言数据稀缺问题
- 在20万亿令牌上优化数据组合与覆盖,实现系统性训练
- 用SnapPO框架高效强化学习,支持多语言推理能力
我们提出Solar Open,一个1020亿参数的双语专家混合语言模型,专为资源匮乏语言设计。针对冷门语言数据不足的问题,我们合成4.5万亿高质量、领域特定且适合强化学习的文本数据。通过渐进式课程机制,在20万亿令牌规模上协同优化数据组成、质量阈值与领域覆盖。为实现可扩展的推理能力,采用自研的SnapPO强化学习框架进行高效优化。在英语与韩语基准测试中,Solar Open表现出色,验证了该方法在冷门语言AI发展中的有效性。
原文摘要 · Abstract (English)
We introduce Solar Open, a 102B-parameter bilingual Mixture-of-Experts language model for underserved languages. Solar Open demonstrates a systematic methodology for building competitive LLMs by addressing three interconnected challenges. First, to train effectively despite data scarcity for underserved languages, we synthesize 4.5T tokens of high-quality, domain-specific, and RL-oriented data. Second, we coordinate this data through a progressive curriculum jointly optimizing composition, quality thresholds, and domain coverage across 20 trillion tokens. Third, to enable reasoning capabilities through scalable RL, we apply our proposed framework SnapPO for efficient optimization. Across benchmarks in English and Korean, Solar Open achieves competitive performance, demonstrating the effectiveness of this methodology for underserved language AI development.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。