25亿参数模型在代码生成上媲美大模型,还能在普通电脑运行。
State-of-the-art Small Language Coder Model: Mify-Coder
- 基于25亿参数模型,用4.2万亿词训练,策略优化计算效率。
- 在编码和函数调用任务中超越更大模型,准确率与安全性俱佳。
- 支持量化部署,普通桌面即可运行,适合实际开发场景。
我们提出 Mify-Coder,一个基于 Mify-2.5B 基础模型、使用 4.2T 标记训练的 2.5B 参数代码模型,采用计算最优策略。Mify-Coder 在标准编码与函数调用基准测试中达到与前沿模型相当的准确率与安全性,显著优于更大规模基线模型,证明小型模型可在代码生成与代理驱动工作流中实现前沿性能。训练流程结合高质量精选数据与通过智能提示生成的合成数据,利用企业级评估数据集迭代优化。基于大模型的质量过滤进一步提升数据密度,实现高效低耗训练。通过系统探索 CPT-SFT 目标、数据混合与采样动态,仅一次连续训练即达成前沿代码智能。实证表明,数据与算力的严谨管理使小模型可获得竞争力强的准确性、效率与安全合规性。量化版本可在标准桌面环境部署,无需专用硬件。
原文摘要 · Abstract (English)
We present Mify-Coder, a 2.5B-parameter code model trained on 4.2T tokens using a compute-optimal strategy built on the Mify-2.5B foundation model. Mify-Coder achieves comparable accuracy and safety while significantly outperforming much larger baseline models on standard coding and function-calling benchmarks, demonstrating that compact models can match frontier-grade models in code generation and agent-driven workflows. Our training pipeline combines high-quality curated sources with synthetic data generated through agentically designed prompts, refined iteratively using enterprise-grade evaluation datasets. LLM-based quality filtering further enhances data density, enabling frugal yet effective training. Through disciplined exploration of CPT-SFT objectives, data mixtures, and sampling dynamics, we deliver frontier-grade code intelligence within a single continuous training trajectory. Empirical evidence shows that principled data and compute discipline allow smaller models to achieve competitive accuracy, efficiency, and safety compliance. Quantized variants of Mify-Coder enable deployment on standard desktop environments without requiring specialized hardware.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。