127亿参数模型通过架构优化实现高效推理,性能媲美更大模型。
Motif 2 12.7B technical report
- 采用分组差异注意力机制分离信号与噪声路径,提升表征效率。
- 在5.5万亿token数据上训练,支持指令泛化与多领域理解。
- 适合资源受限场景下的高效大模型部署,如边缘计算。
我们提出Motif-2-12.7B,一个新发布的开源权重基础模型,通过架构创新与系统级优化推动大语言模型的效率边界。该模型在有限算力预算下具备可扩展的语言理解与强指令泛化能力,基于Motif-2.6B改进,引入分组差异注意力(GDA),通过解耦信号与噪声控制注意力路径提升表征效率。模型在涵盖多种语言、数学、科学和编程领域的5.5万亿标记数据上进行预训练,采用课程驱动的数据调度策略,逐步调整数据组成比例。训练系统结合MuonClip优化器与自定义高性能核函数,包括融合PolyNorm激活和并行Muon算法,在大规模分布式环境中显著提升吞吐量与内存效率。后训练采用三阶段监督微调流程,依次增强通用指令遵循、组合理解与语言精度。Motif-2-12.7B在多个基准测试中表现竞争力,表明精心设计的架构扩展与训练优化可媲美更大规模模型的能力。
原文摘要 · Abstract (English)
We introduce Motif-2-12.7B, a new open-weight foundation model that pushes the efficiency frontier of large language models by combining architectural innovation with system-level optimization. Designed for scalable language understanding and robust instruction generalization under constrained compute budgets, Motif-2-12.7B builds upon Motif-2.6B with the integration of Grouped Differential Attention (GDA), which improves representational efficiency by disentangling signal and noise-control attention pathways. The model is pre-trained on 5.5 trillion tokens spanning diverse linguistic, mathematical, scientific, and programming domains using a curriculum-driven data scheduler that gradually changes the data composition ratio. The training system leverages the MuonClip optimizer alongside custom high-performance kernels, including fused PolyNorm activations and the Parallel Muon algorithm, yielding significant throughput and memory efficiency gains in large-scale distributed environments. Post-training employs a three-stage supervised fine-tuning pipeline that successively enhances general instruction adherence, compositional understanding, and linguistic precision. Motif-2-12.7B demonstrates competitive performance across diverse benchmarks, showing that thoughtful architectural scaling and optimized training design can rival the capabilities of much larger models.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。