arXiv:2512.06266cs.CL2025-12被引 5

3B参数小模型突破性能极限,媲美更大模型。

Nanbeige4-3B Technical Report: Exploring the Frontier of Small Language Models

  • 分阶段训练调度+指令微调,提升小模型表现
  • 推理增强与双偏好蒸馏,复杂任务能力显著提升
  • 适合追求高效推理与高性价比的开发者使用

我们提出Nanbeige4-3B系列小型高性能语言模型。在23万亿高质量语料上预训练,并在超过3000万条多样化指令上进行微调,突破了小模型的缩放定律边界。预训练阶段设计细粒度暖启动-稳定-衰减(FG-WSD)调度器,逐步优化数据混合比例以提升性能;后训练阶段采用思辨生成优化与思维链重构联合机制,显著提升复杂任务表现。在SFT后,利用旗舰推理模型通过双偏好蒸馏(DPD)方法对Nanbeige4-3B进行知识蒸馏,进一步提升性能。最后,采用多阶段强化学习,结合可验证奖励与偏好建模,增强推理与人类对齐能力。大量评估显示,Nanbeige4-3B不仅显著优于同规模模型,更在多项基准上媲美更大模型。模型权重已开源:https://huggingface.co/Nanbeige。

原文摘要 · Abstract (English)

We present Nanbeige4-3B, a family of small-scale but high-performing language models. Pretrained on 23T high-quality tokens and finetuned on over 30 million diverse instructions, we extend the boundary of the scaling law for small language models. In pre-training, we design a Fine-Grained Warmup-Stable-Decay (FG-WSD) training scheduler, which progressively refines data mixtures across stages to boost model performance. In post-training, to improve the quality of the SFT data, we design a joint mechanism that integrates deliberative generation refinement and chain-of-thought reconstruction, yielding substantial gains on complex tasks. Following SFT, we employ our flagship reasoning model to distill Nanbeige4-3B through our proposed Dual Preference Distillation (DPD) method, which leads to further performance gains. Finally, a multi-stage reinforcement learning phase was applied, leveraging verifiable rewards and preference modeling to strengthen abilities on both reasoning and human alignment. Extensive evaluations show that Nanbeige4-3B not only significantly outperforms models of comparable parameter scale but also rivals much larger models across a wide range of benchmarks. The model checkpoints are available at https://huggingface.co/Nanbeige.

小模型推理增强蒸馏强化学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。