Qwen2.5-Coder系列模型在代码生成与修复上表现卓越,超越同尺寸更大模型。
Qwen2.5-Coder Technical Report
- 基于Qwen2.5架构,用超5.5万亿词元数据训练,专注代码能力提升。
- 在10多个代码任务中达顶尖水平,7B模型性能优于更大同类模型。
- 开源许可宽松,适合开发者用于实际编程辅助与研究创新。
本文介绍Qwen2.5-Coder系列,是CodeQwen1.5的重大升级。该系列包含六个模型:Qwen2.5-Coder-(0.5B/1.5B/3B/7B/14B/32B)。作为专用于代码的模型,Qwen2.5-Coder基于Qwen2.5架构,继续在超过5.5万亿个词元的庞大语料上预训练。通过精细的数据清洗、可扩展的合成数据生成和均衡的数据混合策略,该模型在代码生成方面表现出色,同时保持通用性和数学能力。这些模型已在多种代码相关任务上进行评估,在超过10个基准测试中取得领先(SOTA)表现,涵盖代码生成、补全、推理与修复,且始终优于同规模的更大模型。我们相信,Qwen2.5-Coder系列的发布将推动代码智能研究发展,并因其宽松的许可证,支持开发人员在真实应用中广泛采用。
原文摘要 · Abstract (English)
In this report, we introduce the Qwen2.5-Coder series, a significant upgrade from its predecessor, CodeQwen1.5. This series includes six models: Qwen2.5-Coder-(0.5B/1.5B/3B/7B/14B/32B). As a code-specific model, Qwen2.5-Coder is built upon the Qwen2.5 architecture and continues pretrained on a vast corpus of over 5.5 trillion tokens. Through meticulous data cleaning, scalable synthetic data generation, and balanced data mixing, Qwen2.5-Coder demonstrates impressive code generation capabilities while retaining general and math skills. These models have been evaluated on a wide range of code-related tasks, achieving state-of-the-art (SOTA) performance across more than 10 benchmarks, including code generation, completion, reasoning, and repair, consistently outperforming larger models of the same model size. We believe that the release of the Qwen2.5-Coder series will advance research in code intelligence and, with its permissive licensing, support wider adoption by developers in real-world applications.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。