开源高效混合模型,推理速度比同类快3.3倍
Nemotron 3 Nano: Open, Efficient Mixture-of-Experts Hybrid Mamba-Transformer Model for Agentic Reasoning
- 采用Mamba-Transformer混合架构,动态激活少于一半参数
- 在100万词长上下文下保持高准确率,推理吞吐提升3.3倍
- 适合需要高效推理的智能体与复杂任务场景
我们提出Nemotron 3 Nano 30B-A3B,一个基于Mixture-of-Experts的混合Mamba-Transformer语言模型。该模型在25万亿文本标记上预训练,包含超过3万亿新唯一标记,相比Nemotron 2新增。随后经过监督微调和大规模强化学习优化。相比前代Nemotron 2 Nano,Nemotron 3 Nano在每前向传播中激活参数少于一半,仍保持更高准确率。其推理吞吐量最高达同等规模开源模型GPT-OSS-20B和Qwen3-30B-A3B-Thinking-2507的3.3倍,且在多个基准测试中表现更优。模型具备更强的智能体行为、推理与对话能力,支持高达100万标记的上下文长度。我们已在Hugging Face公开发布预训练版Nemotron 3 Nano 30B-A3B Base及后续训练版Nemotron 3 Nano 30B-A3B检查点。
原文摘要 · Abstract (English)
We present Nemotron 3 Nano 30B-A3B, a Mixture-of-Experts hybrid Mamba-Transformer language model. Nemotron 3 Nano was pretrained on 25 trillion text tokens, including more than 3 trillion new unique tokens over Nemotron 2, followed by supervised fine tuning and large-scale RL on diverse environments. Nemotron 3 Nano achieves better accuracy than our previous generation Nemotron 2 Nano while activating less than half of the parameters per forward pass. It achieves up to 3.3x higher inference throughput than similarly-sized open models like GPT-OSS-20B and Qwen3-30B-A3B-Thinking-2507, while also being more accurate on popular benchmarks. Nemotron 3 Nano demonstrates enhanced agentic, reasoning, and chat abilities and supports context lengths up to 1M tokens. We release both our pretrained Nemotron 3 Nano 30B-A3B Base and post-trained Nemotron 3 Nano 30B-A3B checkpoints on Hugging Face.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。