arXiv:2505.17473cs.CVcs.AI2025-05被引 3

构建百万级图文元素检测数据集,提升视觉语言模型对图表的理解能力。

InfoDet: A Dataset for Infographic Element Detection

  • 结合模型与程序化方法生成11万张真实与合成图文
  • 含超1400万边界框标注,覆盖图表与可识别对象
  • 支持图表理解、布局分析等下游任务,推动多模态研究

鉴于图表在科研、商业和传播中的核心作用,提升视觉语言模型(VLMs)的图表理解能力变得日益重要。现有VLMs在图文元素(如图表、图标、图像等)的视觉定位上存在显著误差。为解决这一问题,我们提出InfoDet,一个用于支持图表与人类可识别对象(HROs)精准检测的数据集。该数据集包含11,264张真实图文和90,000张合成图文,涵盖超过1400万条边界框标注,通过模型驱动与程序化方法联合生成。我们展示了InfoDet的实用性:1)构建“基于框推理”框架以提升VLM的图表理解性能;2)对比现有目标检测模型;3)将训练出的检测模型应用于文档版面与UI元素检测任务。

原文摘要 · Abstract (English)

Given the central role of charts in scientific, business, and communication contexts, enhancing the chart understanding capabilities of vision-language models (VLMs) has become increasingly critical. A key limitation of existing VLMs lies in their inaccurate visual grounding of infographic elements, including charts and human-recognizable objects (HROs) such as icons and images. However, chart understanding often requires identifying relevant elements and reasoning over them. To address this limitation, we introduce InfoDet, a dataset designed to support the development of accurate object detection models for charts and HROs in infographics. It contains 11,264 real and 90,000 synthetic infographics, with over 14 million bounding box annotations. These annotations are created by combining the model-in-the-loop and programmatic methods. We demonstrate the usefulness of InfoDet through three applications: 1) constructing a Thinking-with-Boxes scheme to boost the chart understanding performance of VLMs, 2) comparing existing object detection models, and 3) applying the developed detection model to document layout and UI element detection.

图文理解目标检测数据集视觉语言模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。