首个端到端统一文本与视觉信息识别的OCR方法
OCRVerse: Towards Holistic OCR in End-to-End Vision-Language Models
- 构建跨领域数据集,融合文本与图表等视觉密集型图像
- 采用两阶段微调+强化学习实现多场景统一识别
- 支持灵活奖励机制,适配不同输出格式需求
大视觉语言模型的发展推动了对海量多模态数据的管理与应用需求,使从图像中提取信息的OCR技术愈发重要。然而现有OCR方法主要聚焦于从图像或扫描文档中识别文字(以文本为中心),忽视了从图表、网页、科学图示等视觉信息密集型图像中识别视觉元素(以视觉为中心)。这类图像在互联网中广泛存在,在数据可视化和网页分析等领域具有重要应用价值。本文提出OCRVerse,首个以端到端方式实现文本与视觉双重识别的综合型OCR方法。为此,我们构建涵盖报纸、杂志、书籍等文本类文档,以及图表、网页、科学图示等视觉复合图像的全面数据集。同时提出两阶段SFT-RL多域训练策略:SFT阶段通过混合跨域数据建立初始领域知识;RL阶段针对各领域特性设计个性化奖励策略,根据不同任务输出格式要求,灵活定制奖励信号,提升跨域融合能力并避免数据冲突。实验表明,OCRVerse在文本与视觉两类数据上均取得竞争力表现,甚至可媲美大型开源与闭源模型。
原文摘要 · Abstract (English)
The development of large vision language models drives the demand for managing, and applying massive amounts of multimodal data, making OCR technology, which extracts information from visual images, increasingly popular. However, existing OCR methods primarily focus on recognizing text elements from images or scanned documents (Text-centric OCR), neglecting the identification of visual elements from visually information-dense image sources (Vision-centric OCR), such as charts, web pages and science plots. In reality, these visually information-dense images are widespread on the internet and have significant real-world application value, such as data visualization and web page analysis. In this technical report, we propose OCRVerse, the first holistic OCR method in end-to-end manner that enables unified text-centric OCR and vision-centric OCR. To this end, we constructe comprehensive data engineering to cover a wide range of text-centric documents, such as newspapers, magazines and books, as well as vision-centric rendered composites, including charts, web pages and scientific plots. Moreover, we propose a two-stage SFT-RL multi-domain training method for OCRVerse. SFT directly mixes cross-domain data to train and establish initial domain knowledge, while RL focuses on designing personalized reward strategies for the characteristics of each domain. Specifically, since different domains require various output formats and expected outputs, we provide sufficient flexibility in the RL stage to customize flexible reward signals for each domain, thereby improving cross-domain fusion and avoiding data conflicts. Experimental results demonstrate the effectiveness of OCRVerse, achieving competitive results across text-centric and vision-centric data types, even comparable to large-scale open-source and closed-source models.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。