arXiv:2409.01704cs.CV2024-09被引 209

提出通用OCR理论与统一模型GOT,支持多类型文本智能识别

General OCR Theory: Towards OCR-2.0 via a Unified End-to-end Model

论文配图:General OCR Theory: Towards OCR-2.0 via a Unified End-to-end Model
图 1 · 摘自论文原文
  • 构建统一端到端模型GOT,整合多种字符识别任务
  • 580M参数模型支持动态分辨率与跨页识别,输出格式可定制
  • 适用于文档、公式、图表等复杂场景,适合工业级应用

传统OCR系统(OCR-1.0)难以满足人们对人造光学字符智能处理的需求。本文将各类人工光学信号(如文字、数学/分子式、表格、图表、乐谱及几何图形)统称为“字符”,提出通用OCR理论,并构建优秀模型GOT,推动OCR-2.0的到来。GOT拥有580M参数,采用高压缩编码器与长上下文解码器,为统一、优雅的端到端模型。作为OCR-2.0模型,GOT可在多种任务中处理上述所有“字符”。输入支持场景图与整页文档,输出可通过简单提示生成纯文本或格式化结果(如markdown/tikz/smiles/kern)。模型还具备交互式识别功能,支持坐标或颜色引导的区域识别。此外,引入动态分辨率与多页识别技术以增强实用性。实验充分验证了该模型的优越性。

原文摘要 · Abstract (English)

Traditional OCR systems (OCR-1.0) are increasingly unable to meet people's usage due to the growing demand for intelligent processing of man-made optical characters. In this paper, we collectively refer to all artificial optical signals (e.g., plain texts, math/molecular formulas, tables, charts, sheet music, and even geometric shapes) as "characters" and propose the General OCR Theory along with an excellent model, namely GOT, to promote the arrival of OCR-2.0. The GOT, with 580M parameters, is a unified, elegant, and end-to-end model, consisting of a high-compression encoder and a long-contexts decoder. As an OCR-2.0 model, GOT can handle all the above "characters" under various OCR tasks. On the input side, the model supports commonly used scene- and document-style images in slice and whole-page styles. On the output side, GOT can generate plain or formatted results (markdown/tikz/smiles/kern) via an easy prompt. Besides, the model enjoys interactive OCR features, i.e., region-level recognition guided by coordinates or colors. Furthermore, we also adapt dynamic resolution and multi-page OCR technologies to GOT for better practicality. In experiments, we provide sufficient results to prove the superiority of our model.

通用OCR端到端多模态识别GOT模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。