GutenOCR让AI读懂文档,支持精准定位和问答。
GutenOCR: A Grounded Vision-Language Front-End for Documents
- 用提示词统一接口,实现读取、检测与定位一体化。
- 在10500页文档上,接地识别得分从0.40提升至0.82。
- 适合需要精准文本定位的学术与商务文档处理场景。
GutenOCR 是通过微调 Qwen2.5-VL-3B 与 Qwen2.5-VL-7B 构建的一系列接地式 OCR 前端模型。这些单检查点视觉语言模型通过统一的提示接口,实现阅读、检测与接地功能。模型在商业文档、科研文章及合成接地数据上训练,支持整页与局部阅读,提供行级与段落级边界框,并可响应条件性问题如“x在哪里?”。我们提出一种接地式 OCR 评估协议,在10,500张保留的商业与科学页面上,GutenOCR-7B 的综合接地识别得分超过其原始模型(Qwen2.5-VL-7B)的两倍(0.40 → 0.82)。在 Fox 与 OmniDocBench v1.5 上,该方法显著提升区域与行级 OCR 及文本检测召回率,但在页面线性化、颜色引导 OCR 及公式密集布局中出现权衡。
原文摘要 · Abstract (English)
GutenOCR is a family of grounded OCR front-ends obtained by fine-tuning Qwen2.5-VL-3B and Qwen2.5-VL-7B. The resulting single-checkpoint vision-language models expose reading, detection, and grounding through a unified, prompt-based interface. Trained on business documents, scientific articles, and synthetic grounding data, the models support full-page and localized reading with line- and paragraph-level bounding boxes and conditional ``where is x?'' queries. We introduce a grounded OCR evaluation protocol and show that GutenOCR-7B more than doubles the composite grounded OCR score of its Qwen2.5-VL-7B backbone on 10.5K held-out business and scientific pages (0.40 to 0.82). On Fox and OmniDocBench v1.5, our approach substantially improves region- and line-level OCR as well as text-detection recall, but reveals trade-offs in page-level linearization, color-guided OCR, and formula-heavy layouts.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。