10亿参数模型直接从扫描文档生成文本,速度快9倍且精度领先。
LightOnOCR: A 1B End-to-End Multilingual Vision-Language Model for State-of-the-Art OCR
- 端到端训练,跳过传统脆弱的OCR流程
- 在OlmOCR-Bench上达到当前最佳性能,体积小9倍
- 支持图像位置预测,适合需要定位的文档处理场景
我们提出LightOnOCR-2-1B,一个10亿参数的端到端多语言视觉-语言模型,可将文档图像(如PDF)直接转换为清晰、自然排序的文本,无需依赖脆弱的OCR流水线。该模型在大规模高质量蒸馏数据集上训练,覆盖大量扫描件、法语文档及科学类PDF。在OlmOCR-Bench评测中表现领先,相比之前最优模型缩小9倍,且显著更快。进一步扩展输出格式以预测嵌入图像的标准化边界框,通过恢复策略在预训练中引入定位能力,并使用基于IoU奖励的RLVR进行优化。最后通过检查点平均与任务算术合并提升鲁棒性。模型权重已开源(Apache 2.0),数据集及LightOnOCR-bbox-bench评估基准也公开发布。
原文摘要 · Abstract (English)
We present LightOnOCR-2-1B, a 1B-parameter end-to-end multilingual vision--language model that converts document images (e.g., PDFs) into clean, naturally ordered text without brittle OCR pipelines. Trained on a large-scale, high-quality distillation mix with strong coverage of scans, French documents, and scientific PDFs, LightOnOCR-2 achieves state-of-the-art results on OlmOCR-Bench while being 9$\times$ smaller and substantially faster than prior best-performing models. We further extend the output format to predict normalized bounding boxes for embedded images, introducing localization during pretraining via a resume strategy and refining it with RLVR using IoU-based rewards. Finally, we improve robustness with checkpoint averaging and task-arithmetic merging. We release model checkpoints under Apache 2.0, and publicly release the dataset and LightOnOCR-bbox-bench evaluation under their respective licenses.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。