arXiv:2503.02304cs.CV2025-03ICCV被引 14

首个面向图文任务的细粒度文本图像模型,提升文档理解精度

A Token-level Text Image Foundation Model for Document Understanding

  • 构建基于令牌级别的视觉基础模型,精准捕捉密集小文本
  • 训练数据含2000万张图与18亿文本标记对,支持细粒度监督
  • 可替代传统模型,适用于文档问答等实际场景

近年来,通用视觉基础模型(VFMs)在多模态大语言模型中被广泛应用,尤其作为图像编码器。然而,在涉及小而密集文本的下游图文任务(如感知、理解与推理)中,由于缺乏语义细粒度监督,这些模型仍存在根本性预测错误。为此,我们提出TokenOCR,首个专为图文任务设计的令牌级视觉基础模型,可支持多种传统下游应用。为支持TokenOCR的预训练,我们还构建了首个令牌级图文数据集TokenIT,包含2000万张图像和18亿个文本标记-掩码对。进一步地,利用其卓越的图文表征能力,我们以TokenOCR替换原有视觉编码器,构建了面向文档理解的文档级多模态大语言模型TokenVL。大量实验证明了TokenOCR与TokenVL的有效性。代码、数据集及权重将公开于https://github.com/Token-family/TokenFD。

原文摘要 · Abstract (English)

In recent years, general visual foundation models (VFMs) have witnessed increasing adoption, particularly as image encoders for popular multi-modal large language models (MLLMs). However, without semantically fine-grained supervision, these models still encounter fundamental prediction errors in the context of downstream text-image-related tasks, i.e., perception, understanding and reasoning with images containing small and dense texts. To bridge this gap, we develop TokenOCR, the first token-level visual foundation model specifically tailored for text-image-related tasks, designed to support a variety of traditional downstream applications. To facilitate the pretraining of TokenOCR, we also devise a high-quality data production pipeline that constructs the first token-level image text dataset, TokenIT, comprising 20 million images and 1.8 billion token-mask pairs. Furthermore, leveraging this foundation with exceptional image-as-text capability, we seamlessly replace previous VFMs with TokenOCR to construct a document-level MLLM, TokenVL, for VQA-based document understanding tasks. Finally, extensive experiments demonstrate the effectiveness of TokenOCR and TokenVL. Code, datasets, and weights will be available at https://github.com/Token-family/TokenFD.

图文理解视觉模型文档问答细粒度识别

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。