0.9B小模型实现野外文档解析新SOTA,支持多任务且抗物理畸变
PaddleOCR-VL-1.5: Towards a Multi-Task 0.9B VLM for Robust In-the-Wild Document Parsing
- 0.9B参数量的视觉语言模型,统一处理文本识别与多任务解析
- 在真实场景畸变测试集上达94.5%准确率,显著优于现有方法
- 适用于移动端或边缘设备,适合需要轻量化鲁棒文档理解的场景
我们提出PaddleOCR-VL-1.5,一个升级版模型,在OmniDocBench v1.5上达到94.5%的新SOTA准确率。为严格评估真实世界物理畸变(如扫描、倾斜、扭曲、屏幕拍照、光照变化)下的鲁棒性,我们构建了Real5-OmniDocBench基准测试集。实验表明,该模型在新基准上表现优异。此外,模型扩展支持印章识别与文本定位任务,仍保持0.9B超紧凑参数量与高效率。代码已开源。
原文摘要 · Abstract (English)
We introduce PaddleOCR-VL-1.5, an upgraded model achieving a new state-of-the-art (SOTA) accuracy of 94.5% on OmniDocBench v1.5. To rigorously evaluate robustness against real-world physical distortions, including scanning, skew, warping, screen-photography, and illumination, we propose the Real5-OmniDocBench benchmark. Experimental results demonstrate that this enhanced model attains SOTA performance on the newly curated benchmark. Furthermore, we extend the model's capabilities by incorporating seal recognition and text spotting tasks, while remaining a 0.9B ultra-compact VLM with high efficiency. Code: https://github.com/PaddlePaddle/PaddleOCR
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。