NVIDIA 推出 Nemotron 3 系列模型,高效开源,支持超长上下文与智能推理。
NVIDIA Nemotron 3: Efficient and Open Intelligence
- 采用混合专家的 Mamba-Transformer 架构,支持长达 100 万词元上下文。
- Ultra 模型在推理和准确性上达到业界领先水平,Nano 模型性价比突出。
- 适合构建智能代理、自动化任务,开放发布模型与训练全链路资源。
我们推出 Nemotron 3 系列模型——Nano、Super 与 Ultra。这些模型具备强大的代理能力、推理与对话性能。Nemotron 3 系列采用混合专家的 Mamba-Transformer 架构,实现行业领先的吞吐量和高达 100 万词元的上下文长度。Super 与 Ultra 模型使用 NVFP4 训练,并引入新型隐式专家(LatentMoE)技术提升模型质量。两个大模型还集成 MTP 层以加速文本生成。所有模型均通过多环境强化学习后训练,支持推理、多步工具调用及细粒度推理预算控制。Nano 模型在准确率上超越同类模型,且推理成本极低。Super 专为协作代理和高负载场景(如 IT 工单自动化)优化。Ultra 模型在准确率和推理性能上达到当前最优。Nano 模型已随技术报告和白皮书发布,Super 与 Ultra 将陆续开放。我们将公开模型权重、预训练与后训练软件、训练配方及具备分发权限的数据集。
原文摘要 · Abstract (English)
We introduce the Nemotron 3 family of models - Nano, Super, and Ultra. These models deliver strong agentic, reasoning, and conversational capabilities. The Nemotron 3 family uses a Mixture-of-Experts hybrid Mamba-Transformer architecture to provide best-in-class throughput and context lengths of up to 1M tokens. Super and Ultra models are trained with NVFP4 and incorporate LatentMoE, a novel approach that improves model quality. The two larger models also include MTP layers for faster text generation. All Nemotron 3 models are post-trained using multi-environment reinforcement learning enabling reasoning, multi-step tool use, and support granular reasoning budget control. Nano, the smallest model, outperforms comparable models in accuracy while remaining extremely cost-efficient for inference. Super is optimized for collaborative agents and high-volume workloads such as IT ticket automation. Ultra, the largest model, provides state-of-the-art accuracy and reasoning performance. Nano is released together with its technical report and this white paper, while Super and Ultra will follow in the coming months. We will openly release the model weights, pre- and post-training software, recipes, and all data for which we hold redistribution rights.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。