Phi-4用高质量数据超越老师GPT-4,小模型也能强推理
Phi-4 Technical Report
- 用合成数据+高质量训练,不依赖原始网页内容
- 140亿参数在理科问答上超过GPT-4,推理能力突出
- 适合关注高效训练和推理能力的AI研究者
我们介绍phi-4,一个拥有140亿参数的语言模型,其训练策略核心聚焦于数据质量。与多数以网络内容或代码等自然数据为主进行预训练的模型不同,phi-4在训练过程中系统性地引入合成数据。尽管前代Phi系列模型主要通过蒸馏方式模仿教师模型(特别是GPT-4),phi-4在侧重科学、技术、工程与数学的问答任务中显著超越其教师模型,表明我们的数据生成与后训练技术已超出简单蒸馏。尽管对phi-3架构改动极小,凭借优化的数据、训练课程及后训练方案,phi-4在推理类基准测试中表现强劲,相对其规模展现出优异性能。
原文摘要 · Abstract (English)
We present phi-4, a 14-billion parameter language model developed with a training recipe that is centrally focused on data quality. Unlike most language models, where pre-training is based primarily on organic data sources such as web content or code, phi-4 strategically incorporates synthetic data throughout the training process. While previous models in the Phi family largely distill the capabilities of a teacher model (specifically GPT-4), phi-4 substantially surpasses its teacher model on STEM-focused QA capabilities, giving evidence that our data-generation and post-training techniques go beyond distillation. Despite minimal changes to the phi-3 architecture, phi-4 achieves strong performance relative to its size -- especially on reasoning-focused benchmarks -- due to improved data, training curriculum, and innovations in the post-training scheme.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。