Nemotron 3 Ultra 是5500亿参数的高效混合专家模型,支持百万级上下文,适合长期自主任务。
Nemotron 3 Ultra: Open, Efficient Mixture-of-Experts Hybrid Mamba-Transformer Model for Agentic Reasoning

- 采用混合专家与Mamba-Transformer结构,结合隐式专家选择和多标记预测提升效率
- 在100万标记上下文下实现比顶尖模型高6倍的推理吞吐量,精度相当
- 开源全版本模型与训练数据,适合研究长序列智能体与高效推理
我们提出Nemotron 3 Ultra,一个总参数量5500亿、活跃参数550亿的混合专家混合Mamba-注意力语言模型。在20万亿文本标记上预训练,将上下文长度扩展至100万标记,并通过监督微调(SFT)、强化学习(RL)和多教师在线蒸馏(MOPD)进行后训练。该模型融合了隐式专家选择(LatentMoE)、多标记预测(MTP)、NVFP4预训练、多环境强化学习验证与反馈(RLVR)、MOPD及推理预算控制等多项关键技术。相比当前公开的领先大模型,其推理吞吐量最高提升约6倍,同时保持相当的准确率。卓越的准确性、高推理吞吐量与100万标记上下文长度,使其特别适用于长时间运行的自主智能体任务。我们已在HuggingFace上开源基础模型、后训练模型及量化版本的检查点、训练数据与完整训练配方。
原文摘要 · Abstract (English)
We introduce Nemotron 3 Ultra, a 550 billion total and 55 billion active parameter Mixture-of-Experts Hybrid Mamba-Attention language model. We pre-trained Nemotron 3 Ultra on 20 trillion text tokens, then extended the context length to 1M tokens, and post-trained using Supervised Fine Tuning (SFT), Reinforcement Learning (RL), and Multi-teacher On-Policy Distillation (MOPD). Nemotron 3 Ultra is our most capable model yet, employing multiple key technologies - LatentMoE, Multi Token Prediction (MTP), NVFP4 pre-training, multi-environment RLVR, MOPD, and reasoning budget control. Nemotron 3 Ultra achieves up to ~6x higher inference throughput as compared to state-of-the-art publicly available LLMs while attaining on-par accuracy. The state-of-the-art accuracy, high inference throughput, and 1M token context length make Nemotron 3 Ultra ideal for long-running autonomous agentic tasks. We open-source the base, post-trained, and quantized checkpoints, along with the training data and recipe on HuggingFace.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。