arXiv:2604.24954cs.LGcs.AI2026-04被引 8

Nemotron 3 Nano Omni 支持多模态输入,推理更快更高效。

Nemotron 3 Nano Omni: Efficient and Open Multimodal Intelligence

论文配图:Nemotron 3 Nano Omni: Efficient and Open Multimodal Intelligence
图 1 · 摘自论文原文
  • 基于30B参数骨架,融合音频、文本、图像与视频的多模态理解
  • 在文档理解、长时音视频分析等任务上超越前代模型,延迟更低
  • 适合需要低延迟多模态推理的研究者与开发者使用

我们推出Nemotron 3 Nano Omni,是Nemotron多模态系列的最新成员,首个原生支持音频输入的模型,兼容文本、图像与视频。该模型在所有模态上均显著优于前代Nemotron Nano V2 VL,得益于架构改进、训练数据与训练策略优化。尤其在真实场景文档理解、长时音视频理解及代理式计算机操作任务中表现领先。基于高效的Nemotron 3 Nano 30B-A3B骨干网络,并引入创新的多模态令牌压缩技术,实现远低于同规模模型的推理延迟与更高吞吐量。我们发布BF16、FP8和FP4格式的模型权重,以及部分训练数据与代码库,以推动后续研究与开发。

原文摘要 · Abstract (English)

We introduce Nemotron 3 Nano Omni, the latest model in the Nemotron multimodal series and the first to natively support audio inputs alongside text, images, and video. Nemotron 3 Nano Omni delivers consistent accuracy improvements over its predecessor, Nemotron Nano V2 VL, across all modalities, enabled by advances in architecture, training data and recipes. In particular, Nemotron 3 delivers leading results in real-world document understanding, long audio-video comprehension, and agentic computer use. Built on the highly efficient Nemotron 3 Nano 30B-A3B backbone, Nemotron 3 Nano Omni further incorporates innovative multimodal token-reduction techniques to deliver substantially lower inference latency and higher throughput than other models of similar size. We are releasing model checkpoints in BF16, FP8, and FP4 formats, along with portions of the training data and codebase to facilitate further research and development.

多模态高效推理音频理解

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。