arXiv:2602.19188cs.CV2026-02

用专家模型提升大模型的定位能力,实现精准文本识别。

PositionOCR: Augmenting Positional Awareness in Multi-Modal Models via Hybrid Specialist Integration

  • 融合文本定位专家与大语言模型,兼顾位置精度与语义理解。
  • 仅131M参数,性能超越传统多模态大模型。
  • 适合需要高精度文本定位的应用,如文档分析、视觉问答。

近年来,多模态大语言模型(MLLM)在以光学字符识别(OCR)为核心的视觉问答任务中表现出色,展现出处理异构数据和跨场景适应的能力。然而,这些模型依赖以语言处理为主的大型语言模型(LLM)作为解码器,天然缺乏精确视觉任务所需的定位推理能力,例如文本定位和文本接地。同时,MLLM庞大的参数量导致训练需大量计算资源和数据。相反,文本定位专家虽能实现最先进的坐标预测,却缺乏语义推理能力。为此,我们提出核心研究问题:能否将专家模型的高效性与大语言模型的上下文理解力结合,构建具有精准定位能力的多模态大模型?为此,我们提出PositionOCR,一种参数高效的混合架构,无缝整合文本定位模型的位置优势与大语言模型的上下文推理能力。该框架仅含131M可训练参数,在文本接地和文本定位等任务上表现卓越,持续优于传统MLLM。

原文摘要 · Abstract (English)

In recent years, Multi-modal Large Language Models (MLLMs) have achieved strong performance in OCR-centric Visual Question Answering (VQA) tasks, illustrating their capability to process heterogeneous data and exhibit adaptability across varied contexts. However, these MLLMs rely on a Large Language Model (LLM) as the decoder, which is primarily designed for linguistic processing, and thus inherently lacks the positional reasoning required for precise visual tasks, such as text spotting and text grounding. Additionally, the extensive parameters of MLLMs necessitate substantial computational resources and large-scale data for effective training. Conversely, text spotting specialists achieve state-of-the-art coordinate predictions but lack semantic reasoning capabilities. This dichotomy motivates our key research question: Can we synergize the efficiency of specialists with the contextual power of LLMs to create a positionally-accurate MLLM? To overcome these challenges, we introduce PositionOCR, a parameter-efficient hybrid architecture that seamlessly integrates a text spotting model's positional strengths with an LLM's contextual reasoning. Comprising 131M trainable parameters, this framework demonstrates outstanding multi-modal processing capabilities, particularly excelling in tasks such as text grounding and text spotting, consistently surpassing traditional MLLMs.

多模态文本定位轻量化大模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。