arXiv:2503.10596cs.CV2025-03ICCV被引 9

构建大规模像素指代数据集,提升视觉语言对齐精度。

GroundingSuite: Measuring Complex Multi-Granular Pixel Grounding

  • 用多模型协同自动标注,生成956万条指代表达与分割图。
  • 在gRefCOCO上达68.9 cIoU,RefCOCOm上达55.3 gIoU。
  • 标注效率比现有方法快4.5倍,适合视觉语言研究者使用。

像素指代任务(如指代表达分割)因有望弥合视觉与语言模态鸿沟而备受关注。然而当前进展受限于现有数据集的不足,包括物体类别有限、文本多样性不足及高质量标注稀缺。为此,我们提出GroundingSuite,包含:(1) 基于多个视觉-语言模型(VLM)代理的自动化数据标注框架;(2) 包含956万条多样化指代表达及其对应分割结果的大规模训练数据集;(3) 精心构建的3,800张图像评估基准。该训练数据集显著提升模型性能,使基于它的模型在gRefCOCO上达到68.9 cIoU,RefCOCOm上达到55.3 gIoU。此外,其标注框架效率优于当前领先方法GLaMM,提速达4.5倍。

原文摘要 · Abstract (English)

Pixel grounding, encompassing tasks such as Referring Expression Segmentation (RES), has garnered considerable attention due to its immense potential for bridging the gap between vision and language modalities. However, advancements in this domain are currently constrained by limitations inherent in existing datasets, including limited object categories, insufficient textual diversity, and a scarcity of high-quality annotations. To mitigate these limitations, we introduce GroundingSuite, which comprises: (1) an automated data annotation framework leveraging multiple Vision-Language Model (VLM) agents; (2) a large-scale training dataset encompassing 9.56 million diverse referring expressions and their corresponding segmentations; and (3) a meticulously curated evaluation benchmark consisting of 3,800 images. The GroundingSuite training dataset facilitates substantial performance improvements, enabling models trained on it to achieve state-of-the-art results. Specifically, a cIoU of 68.9 on gRefCOCO and a gIoU of 55.3 on RefCOCOm. Moreover, the GroundingSuite annotation framework demonstrates superior efficiency compared to the current leading data annotation method, i.e., $4.5 \times$ faster than GLaMM.

像素指代多模态数据集视觉语言

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。