arXiv:2511.20478cs.LG2025-11

轻量级文档解析模型,支持多类型内容高效提取

NVIDIA Nemotron Parse 1.1

  • 采用编码器-解码器架构,885M参数,语言解码器256M
  • 在公开基准上表现优异,支持长序列输出与复杂图文解析
  • 提供加速版模型,速度提升20%且质量损失小,适合部署

我们提出Nemotron-Parse-1.1,一种轻量级文档解析与OCR模型,相较于前代Nemoretriever-Parse-1.0有全面升级。该模型在通用OCR、Markdown格式化、结构化表格解析及图片、图表和示意图中的文本提取方面表现更优,并支持对视觉密集型文档的更长输出序列。与前代一致,可提取文本段落的边界框及其语义类别。模型采用编码器-解码器架构,共885M参数,其中语言解码器为256M参数。在公共基准上达到竞争性准确率,是优秀的轻量级OCR解决方案。我们已在Huggingface公开模型权重,同时提供优化的NIM容器及部分训练数据(作为Nemotron-VLM-v2数据集的一部分)。此外,还发布Nemotron-Parse-1.1-TC版本,通过减少视觉标记长度实现20%的速度提升,质量损失极小。

原文摘要 · Abstract (English)

We introduce Nemotron-Parse-1.1, a lightweight document parsing and OCR model that advances the capabilities of its predecessor, Nemoretriever-Parse-1.0. Nemotron-Parse-1.1 delivers improved capabilities across general OCR, markdown formatting, structured table parsing, and text extraction from pictures, charts, and diagrams. It also supports a longer output sequence length for visually dense documents. As with its predecessor, it extracts bounding boxes of text segments, as well as corresponding semantic classes. Nemotron-Parse-1.1 follows an encoder-decoder architecture with 885M parameters, including a compact 256M-parameter language decoder. It achieves competitive accuracy on public benchmarks making it a strong lightweight OCR solution. We release the model weights publicly on Huggingface, as well as an optimized NIM container, along with a subset of the training data as part of the broader Nemotron-VLM-v2 dataset. Additionally, we release Nemotron-Parse-1.1-TC which operates on a reduced vision token length, offering a 20% speed improvement with minimal quality degradation.

文档解析轻量模型OCRNVIDIA

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。