开源20亿参数模型,高效训练下性能媲美顶尖模型
PCMind-2.1-Kaiyuan-2B Technical Report
- 用分位数数据基准评估开源数据集,优化混合策略
- 通过多阶段选择性重复,高效利用稀缺高质量数据
- 按质量排序的多领域课程训练,提升资源受限下的效果
大型语言模型的快速发展导致开源社区与工业界之间存在显著知识鸿沟,主要源于后者依赖封闭的高质量数据和训练方法。为此,我们推出全开源的20亿参数模型PCMind-2.1-Kaiyuan-2B,专注于在资源受限条件下提升训练效率与效果。方法包括:基于分位数的数据基准评估法,系统比较异构开源数据集并提供数据混合策略;多阶段范式中的战略性选择性重复机制,有效利用稀疏高质量数据;按质量排序的多领域课程训练策略。结合高度优化的数据预处理流水线及支持FP16稳定性的架构改进,该模型在性能上达到当前全开源领先水平,为资源有限场景下的可扩展预训练提供了实用方案。所有资产(含模型权重、数据、代码)均以Apache 2.0协议开源,地址为https://huggingface.co/thu-pacman/PCMind-2.1-Kaiyuan-2B。
原文摘要 · Abstract (English)
The rapid advancement of Large Language Models (LLMs) has resulted in a significant knowledge gap between the open-source community and industry, primarily because the latter relies on closed-source, high-quality data and training recipes. To address this, we introduce PCMind-2.1-Kaiyuan-2B, a fully open-source 2-billion-parameter model focused on improving training efficiency and effectiveness under resource constraints. Our methodology includes three key innovations: a Quantile Data Benchmarking method for systematically comparing heterogeneous open-source datasets and providing insights on data mixing strategies; a Strategic Selective Repetition scheme within a multi-phase paradigm to effectively leverage sparse, high-quality data; and a Multi-Domain Curriculum Training policy that orders samples by quality. Supported by a highly optimized data preprocessing pipeline and architectural modifications for FP16 stability, Kaiyuan-2B achieves performance competitive with state-of-the-art fully open-source models, demonstrating practical and scalable solutions for resource-limited pretraining. We release all assets (including model weights, data, and code) under Apache 2.0 license at https://huggingface.co/thu-pacman/PCMind-2.1-Kaiyuan-2B.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。