arXiv:2603.13398cs.CV2026-03被引 17

40亿参数模型一键解析文档,支持多种任务并精准还原布局。

Qianfan-OCR: A Unified End-to-End Model for Document Intelligence

  • 统一架构实现图像直接转Markdown,支持表格、图表等多任务
  • 引入布局思考机制,在复杂文档中准确生成元素位置与顺序
  • 在多个公开评测中超越Gemini等大模型,适合文档智能场景

我们提出Qianfan-OCR,一个40亿参数的端到端视觉语言模型,将文档解析、版面分析和文档理解统一于单一架构中。该模型可直接实现图像到Markdown的转换,并支持表格提取、图表理解、文档问答和关键信息抽取等多种提示驱动任务。为解决端到端OCR中版面分析缺失的问题,我们提出Layout-as-Thought机制——通过特殊思考令牌触发,生成包含边界框、元素类型和阅读顺序的结构化布局表示,从而恢复布局感知能力并提升复杂版面的准确性。Qianfan-OCR在OmniDocBench v1.5(93.12)和OlmOCR Bench(79.8)上位列所有端到端模型第一,在OCRBench、CCOCR、DocVQA和ChartQA上表现优于同规模通用视觉语言模型,并在公开的关键信息抽取基准上取得最高平均分,超过Gemini-3.1-Pro、Seed-2.0和Qwen3-VL-235B。模型可通过百度智能云千帆平台免费获取。

原文摘要 · Abstract (English)

We present Qianfan-OCR, a 4B-parameter end-to-end vision-language model that unifies document parsing, layout analysis, and document understanding within a single architecture. It performs direct image-to-Markdown conversion and supports diverse prompt-driven tasks including table extraction, chart understanding, document QA, and key information extraction. To address the loss of explicit layout analysis in end-to-end OCR, we propose Layout-as-Thought, an optional thinking phase triggered by special think tokens that generates structured layout representations -- bounding boxes, element types, and reading order -- before producing final outputs, recovering layout grounding capabilities while improving accuracy on complex layouts. Qianfan-OCR ranks first among end-to-end models on OmniDocBench v1.5 (93.12) and OlmOCR Bench (79.8), achieves competitive results on OCRBench, CCOCR, DocVQA, and ChartQA against general VLMs of comparable scale, and attains the highest average score on public key information extraction benchmarks, surpassing Gemini-3.1-Pro, Seed-2.0, and Qwen3-VL-235B. The model is publicly accessible via the Baidu AI Cloud Qianfan platform.

文档理解视觉语言模型端到端OCR布局分析

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。