arXiv:2507.05595cs.CV2025-07被引 152

PaddleOCR 3.0用小模型实现高效多语言文档理解

PaddleOCR 3.0 Technical Report

  • 采用轻量级模型实现多语言文字识别与文档结构解析
  • 百万元级参数模型达到千万元级大模型的识别精度
  • 支持跨硬件部署,适合开发智能文档应用的开发者

本技术报告介绍 PaddleOCR 3.0,一个 Apache 许可的开源工具包,用于 OCR 和文档解析。为应对大语言模型时代对文档理解日益增长的需求,PaddleOCR 3.0 提出三大解决方案:(1) PP-OCRv5 实现多语言文本识别,(2) PP-StructureV3 支持层次化文档解析,(3) PP-ChatOCRv4 用于关键信息抽取。相比主流视觉语言模型(VLMs),这些模型参数少于 1 亿,在准确率和效率上表现优异,媲美数十亿参数的 VLMs。除提供高质量的 OCR 模型库外,还配备高效的训练、推理与部署工具,支持异构硬件加速,帮助开发者便捷构建智能文档应用。

原文摘要 · Abstract (English)

This technical report introduces PaddleOCR 3.0, an Apache-licensed open-source toolkit for OCR and document parsing. To address the growing demand for document understanding in the era of large language models, PaddleOCR 3.0 presents three major solutions: (1) PP-OCRv5 for multilingual text recognition, (2) PP-StructureV3 for hierarchical document parsing, and (3) PP-ChatOCRv4 for key information extraction. Compared to mainstream vision-language models (VLMs), these models with fewer than 100 million parameters achieve competitive accuracy and efficiency, rivaling billion-parameter VLMs. In addition to offering a high-quality OCR model library, PaddleOCR 3.0 provides efficient tools for training, inference, and deployment, supports heterogeneous hardware acceleration, and enables developers to easily build intelligent document applications.

OCR文档解析轻量化模型中文识别

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。