arXiv:2504.04974cs.CVcs.AI2025-04被引 18

针对文档图像中的文本定位难题,提出新任务TRIG与数据集。

Towards Visual Text Grounding of Multimodal Large Language Model

  • 构建OCR-LLM-human协同标注流程,生成800人工标注对。
  • 设计90万条合成数据训练集,显著提升模型在复杂文档上的定位能力。
  • 适合研究多模态模型视觉文本对齐的开发者和研究员。

尽管多模态大语言模型(MLLMs)不断发展,其在文本密集型图像(如扫描表单、信息图)中的视觉文本定位能力仍存在明显不足。现有基准主要聚焦自然图像,未能充分覆盖文档图像的复杂版式与丰富文本。为此,我们提出TRIG任务,并构建新的指令数据集以评估和改进MLLM在文档问答中的文本定位能力。通过OCR-LLM-人类协作流程,创建800个手工标注的问答对作为基准与小规模训练集,同时基于四个不同数据集生成90万条合成数据用于大规模训练。对多种MLLM在该基准上的全面评估揭示了其在文本密集图像上的显著局限性。此外,我们提出两种简单有效的TRIG方法:基于通用指令微调与即插即用高效嵌入。在合成数据上微调后,模型的空间推理与定位能力得到显著提升。

原文摘要 · Abstract (English)

Despite the existing evolution of Multimodal Large Language Models (MLLMs), a non-neglectable limitation remains in their struggle with visual text grounding, especially in text-rich images of documents. Document images, such as scanned forms and infographics, highlight critical challenges due to their complex layouts and textual content. However, current benchmarks do not fully address these challenges, as they mostly focus on visual grounding on natural images, rather than text-rich document images. Thus, to bridge this gap, we introduce TRIG, a novel task with a newly designed instruction dataset for benchmarking and improving the Text-Rich Image Grounding capabilities of MLLMs in document question-answering. Specifically, we propose an OCR-LLM-human interaction pipeline to create 800 manually annotated question-answer pairs as a benchmark and a large-scale training set of 90$ synthetic data based on four diverse datasets. A comprehensive evaluation of various MLLMs on our proposed benchmark exposes substantial limitations in their grounding capability on text-rich images. In addition, we propose two simple and effective TRIG methods based on general instruction tuning and plug-and-play efficient embedding, respectively. By finetuning MLLMs on our synthetic dataset, they promisingly improve spatial reasoning and grounding capabilities.

多模态文本定位文档理解大模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。