arXiv:2411.17125cs.CVcs.AI2024-11ICCV被引 11

构建细粒度文档定位与指代数据集,提升多模态模型对文档的精准理解与交互能力。

DOGR: Towards Versatile Visual Document Grounding and Referring

  • 自动生成多粒度文档解析与指令微调数据,增强模型定位与识别能力。
  • 建立覆盖7类任务的DOGR-Bench基准,涵盖图表、海报和PDF三类文档。
  • 提出DOGR模型,实现对话中精准指代关键文本信息,支持灵活交互。

随着多模态大语言模型(MLLMs)的发展,文档中的定位与指代能力日益受到关注,有助于实现细致理解与灵活用户交互。然而,由于细粒度数据集稀缺且缺乏全面评测基准,该能力在视觉文档理解领域仍不成熟。为此,我们提出文档定位与指代数据生成引擎(DOGR-Engine),可生成两类高质量细粒度文档数据:(1) 多粒度解析数据,用于提升文本定位与识别能力;(2) 指令微调数据,激活MLLM在对话与推理中进行定位与指代的能力。基于此,我们构建了DOGR-Bench,涵盖三种文档类型(图表、海报、PDF)下的七项定位与指代任务,提供全面评估。进一步利用生成数据,我们开发出强基线模型DOGR,其在文本定位与识别上表现优异,并能在对话与推理中精准定位与指代关键文本信息,推动文档理解向更细粒度发展,支持灵活交互范式。

原文摘要 · Abstract (English)

With recent advances in Multimodal Large Language Models (MLLMs), grounding and referring capabilities have gained increasing attention for achieving detailed understanding and flexible user interaction. However, these capabilities still remain underdeveloped in visual document understanding due to the scarcity of fine-grained datasets and comprehensive benchmarks. To fill this gap, we propose the DOcument Grounding and Referring data engine (DOGR-Engine), which generates two types of high-quality fine-grained document data: (1) multi-granular parsing data to improve text localization and recognition, and (2) instruction-tuning data to activate MLLMs' grounding and referring capabilities in dialogue and reasoning. Using the DOGR-Engine, we construct DOGR-Bench, a benchmark covering seven grounding and referring tasks across three document types (chart, poster, and PDF document), offering a comprehensive evaluation of fine-grained document understanding. Leveraging the generated data, we further develop DOGR, a strong baseline model that excels in text localization and recognition, while precisely grounds and refers to key textual information during conversation and reasoning, thereby advancing document understanding to a finer granularity and enable flexible interaction paradigms.

文档理解多模态指代识别基准测试

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。