NVIDIA发布新视觉语言模型,提升文档与长视频理解能力
NVIDIA Nemotron Nano V2 VL
- 基于Mamba-Transformer混合架构,结合创新令牌压缩技术
- 相比前代模型在图文任务中全面升级,推理吞吐量显著提升
- 适合需要高效处理长文档和视频的工业级应用
我们推出Nemotron Nano V2 VL,Nemotron视觉语言系列最新模型,专为强现实场景下的文档理解、长视频分析与推理任务设计。该模型在模型架构、数据集和训练方法上实现重大优化,相较于前代Llama-3.1-Nemotron-Nano-VL-8B,在所有视觉与文本领域均取得显著性能提升。Nemotron Nano V2 VL基于Nemotron Nano V2(一种混合Mamba-Transformer大模型),并引入创新的令牌压缩技术,有效提升长文档与视频场景下的推理吞吐量。我们以BF16、FP8和FP4格式发布模型检查点,并开放部分数据集、训练配方与代码。
原文摘要 · Abstract (English)
We introduce Nemotron Nano V2 VL, the latest model of the Nemotron vision-language series designed for strong real-world document understanding, long video comprehension, and reasoning tasks. Nemotron Nano V2 VL delivers significant improvements over our previous model, Llama-3.1-Nemotron-Nano-VL-8B, across all vision and text domains through major enhancements in model architecture, datasets, and training recipes. Nemotron Nano V2 VL builds on Nemotron Nano V2, a hybrid Mamba-Transformer LLM, and innovative token reduction techniques to achieve higher inference throughput in long document and video scenarios. We are releasing model checkpoints in BF16, FP8, and FP4 formats and sharing large parts of our datasets, recipes and training code.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。