arXiv:2501.15558cs.CV2025-01被引 32

3B模型Ocean-OCR在多场景文字识别上超越专业工具,通用理解能力也出色。

Ocean-OCR: Towards General OCR Application via a Vision-Language Model

  • 用原生分辨率视觉编码器支持可变输入尺寸,提升文字识别精度。
  • 在文档、场景、手写文字等6个任务中达到领先水平,部分超专业模型。
  • 首个在多个场景超越TextIn、PaddleOCR的多模态大模型,适合通用文本处理。

多模态大语言模型(MLLM)在跨领域任务中表现出色,但在文字识别方面仍存在不足。本文提出30亿参数的Ocean-OCR模型,在多种文字识别场景下表现优异,同时具备与主流模型相当的通用理解能力。该模型采用原生分辨率视觉变换器(Native Resolution ViT)以支持可变分辨率输入,并利用大规模高质量OCR数据集进行训练。在开源OCR基准和多个实际场景(包括文档理解、场景文本识别、手写体识别)上的全面实验表明,Ocean-OCR性能领先,是首个在多项任务中超越TextIn和PaddleOCR等专业OCR模型的MLLM。

原文摘要 · Abstract (English)

Multimodal large language models (MLLMs) have shown impressive capabilities across various domains, excelling in processing and understanding information from multiple modalities. Despite the rapid progress made previously, insufficient OCR ability hinders MLLMs from excelling in text-related tasks. In this paper, we present \textbf{Ocean-OCR}, a 3B MLLM with state-of-the-art performance on various OCR scenarios and comparable understanding ability on general tasks. We employ Native Resolution ViT to enable variable resolution input and utilize a substantial collection of high-quality OCR datasets to enhance the model performance. We demonstrate the superiority of Ocean-OCR through comprehensive experiments on open-source OCR benchmarks and across various OCR scenarios. These scenarios encompass document understanding, scene text recognition, and handwritten recognition, highlighting the robust OCR capabilities of Ocean-OCR. Note that Ocean-OCR is the first MLLM to outperform professional OCR models such as TextIn and PaddleOCR.

OCR多模态视觉语言模型文本识别

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。