Qwen2.5通过18万亿token训练和多阶段优化,实现更强语言理解与生成能力。
Qwen2.5 Technical Report

- 基于18万亿高质量语料预训练,提升常识与推理基础
- 百万级监督微调+多阶段强化学习,显著改善长文本与指令遵循
- 支持多种规模模型,适合通用任务及垂直领域应用
本文介绍Qwen2.5系列大语言模型,该系列在预训练和后训练阶段均显著优化。预训练阶段,高质量数据集从此前的7万亿扩展至18万亿令牌,为常识、专业知识和推理能力奠定坚实基础。后训练阶段,采用超百万样本的精细监督微调及多阶段强化学习,有效增强人类偏好对齐,显著提升长文本生成、结构化数据分析与指令遵循能力。模型系列覆盖多种尺寸,开放权重版本包括基础与指令微调模型,含量化版本;托管方案提供两款混合专家(MoE)模型:Qwen2.5-Turbo与Qwen2.5-Plus,均可在阿里云模型实验室获取。Qwen2.5在多项基准测试中表现卓越,其开源旗舰模型Qwen2.5-72B-Instruct性能超越多个开源与专有模型,媲美约五倍更大的Llama-3-405B-Instruct。Qwen2.5-Turbo与Qwen2.5-Plus在成本效益上优于GPT-4o-mini与GPT-4o。此外,Qwen2.5作为基座模型,已用于训练Qwen2.5-Math、Qwen2.5-Coder、QwQ及多模态模型。
原文摘要 · Abstract (English)
In this report, we introduce Qwen2.5, a comprehensive series of large language models (LLMs) designed to meet diverse needs. Compared to previous iterations, Qwen 2.5 has been significantly improved during both the pre-training and post-training stages. In terms of pre-training, we have scaled the high-quality pre-training datasets from the previous 7 trillion tokens to 18 trillion tokens. This provides a strong foundation for common sense, expert knowledge, and reasoning capabilities. In terms of post-training, we implement intricate supervised finetuning with over 1 million samples, as well as multistage reinforcement learning. Post-training techniques enhance human preference, and notably improve long text generation, structural data analysis, and instruction following. To handle diverse and varied use cases effectively, we present Qwen2.5 LLM series in rich sizes. Open-weight offerings include base and instruction-tuned models, with quantized versions available. In addition, for hosted solutions, the proprietary models currently include two mixture-of-experts (MoE) variants: Qwen2.5-Turbo and Qwen2.5-Plus, both available from Alibaba Cloud Model Studio. Qwen2.5 has demonstrated top-tier performance on a wide range of benchmarks evaluating language understanding, reasoning, mathematics, coding, human preference alignment, etc. Specifically, the open-weight flagship Qwen2.5-72B-Instruct outperforms a number of open and proprietary models and demonstrates competitive performance to the state-of-the-art open-weight model, Llama-3-405B-Instruct, which is around 5 times larger. Qwen2.5-Turbo and Qwen2.5-Plus offer superior cost-effectiveness while performing competitively against GPT-4o-mini and GPT-4o respectively. Additionally, as the foundation, Qwen2.5 models have been instrumental in training specialized models such as Qwen2.5-Math, Qwen2.5-Coder, QwQ, and multimodal models.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。