12亿参数模型专攻推理,高效训练达顶尖水平
Xmodel-2 Technical Report
- 统一超参设计支持多尺度模型实验与迁移
- 用1.5万亿词预训练,在推理任务上表现最优
- 适合需要低成本高效推理模型的研究者使用
Xmodel-2 是一个专为推理任务设计的12亿参数大语言模型。其架构支持不同规模模型共享统一超参数,可在小模型上广泛实验,并将最优配置无缝迁移至大模型。为提升训练效率与稳定性,Xmodel-2 采用 MiniCPM 的 WSD 学习率调度器。模型在来自多样化来源的1.5万亿词数据上进行预训练,在复杂推理与基于代理的任务中达到当前最优性能,同时保持较低训练成本。这些结果展示了高效模型设计与训练策略在提升推理能力方面的潜力。模型检查点与代码已公开于 GitHub:https://github.com/XiaoduoAILab/Xmodel-2。
原文摘要 · Abstract (English)
Xmodel-2 is a 1.2-billion-parameter large language model designed specifically for reasoning tasks. Its architecture enables different model scales to share a unified set of hyperparameters, allowing for extensive experimentation on smaller models and seamless transfer of optimal configurations to larger models. To maximize training efficiency and stability, Xmodel-2 employs the WSD learning rate scheduler from MiniCPM. Pretrained on 1.5 trillion tokens from diverse sources, Xmodel-2 achieves state-of-the-art performance in complex reasoning and agent-based tasks, while maintaining low training costs. These results highlight the potential of efficient model design and training strategies in advancing reasoning capabilities. Model checkpoints and code are publicly available on GitHub at https://github.com/XiaoduoAILab/Xmodel-2
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。