arXiv:2508.10948cs.LGcs.AI2025-08被引 1

150亿参数模型实现320亿级性能,内存减半

Apriel-Nemotron-15B-Thinker

  • 分四阶段训练:基座扩展+持续预训练+监督微调+强化学习
  • 150亿参数模型在多个基准上媲美320亿参数模型
  • 适合对推理性能与资源消耗敏感的企业场景

尽管大型语言模型在代码、数学等企业任务中展现出卓越的推理能力,但其巨大的内存和计算开销常使其难以在实际企业环境中部署。为此,我们提出 Apriel-Nemotron-15B-Thinker,这是 ServiceNow Apriel SLM 系列中的一个 150 亿参数模型,在多个基准测试中表现可媲美 o1-mini、QWQ32B、EXAONE-Deep-32B 等中型先进模型,同时内存占用仅为这些模型的一半。该模型采用四阶段训练流程:1)基座模型扩增;2)持续预训练;3)监督微调(SFT);4)基于 GRPO 的强化学习。综合评估表明,尽管参数量不足其一半,本模型在各类基准测试中仍能匹配或超越 320 亿参数的同类模型。

原文摘要 · Abstract (English)

While large language models (LLMs) have achieved remarkable reasoning capabilities across domains like code, math and other enterprise tasks, their significant memory and computational costs often preclude their use in practical enterprise settings. To this end, we introduce Apriel-Nemotron-15B-Thinker, a 15-billion parameter model in the ServiceNow Apriel SLM series that achieves performance against medium sized state-of-the-art models such as o1-mini, QWQ32B, and EXAONE-Deep-32B while maintaining only half the memory footprint of those alternatives. Apriel-Nemotron-15B-Thinker model is trained in a four stage training pipeline including 1) Base Model upscaling, 2) Continual Pre-training 3) Supervised Fine-tuning (SFT) and 4) Reinforcement Learning using GRPO. Comprehensive evaluations across a diverse suite of benchmarks consistently demonstrate that our Apriel-Nemotron-15B-Thinker model matches or exceeds the performance of its 32-billion parameter counterparts, despite being less than half their size.

大模型压缩推理优化企业应用

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。