arXiv:2506.24102cs.CV2025-06被引 17

构建首个真实世界细粒度密集视觉定位描述数据集

DenseWorld-1M: Towards Detailed Dense Grounded Caption in the Real World

  • 三阶段标注流程:感知、细粒度描述生成、空间关系合并
  • 涵盖100万条密集描述,支持高分辨率图像中物体位置与关系定位
  • 适用于视觉理解、定位与区域描述任务,适合多模态模型训练

多模态大语言模型在场景理解方面表现复杂,得益于大规模高质量数据集。现有字幕数据集普遍缺乏视觉实体的定位信息和关系描述。部分带定位的数据集存在细节缺失、关系不全及高分辨率图像上物体描述不足的问题。为填补这一空白,我们提出DenseWorld-1M,首个真实世界中大规模、细粒度、密集定位的字幕数据集。设计了三阶段标注流程:开放世界感知、物体级详细描述生成、密集字幕合并。第一阶段获取实体级掩码与标签;第二阶段在掩码和标签引导下生成物体级详细描述;第三阶段将描述与掩码融合为包含空间与关系的密集描述。为加速标注并提升质量,引入两个视觉语言模型:详细区域字幕模型与空间字幕合并模型。在多种设置下的实验(包括视觉语言理解、视觉定位、区域描述生成)验证了DenseWorld-1M数据集与标注模型的有效性。

原文摘要 · Abstract (English)

Multimodal Large Language Models (MLLMs) demonstrate a complex understanding of scenes, benefiting from large-scale and high-quality datasets. Most existing caption datasets lack the ground locations and relations for visual entities. Several grounded caption datasets face the problems of missing detailed descriptions, relations, and massive object descriptions on high-resolution images. To fill this gap for the community, we present DenseWorld-1M, the first massive, detailed, dense grounded caption dataset in the real world. We design a three-stage labeling pipeline, containing open-world perception, detailed object caption generation, and dense caption merging. The first stage obtains entity-level masks and labels. The second stage generates the object-level, detailed captions with the guidance of masks and labels from the first stage. The final stage merges object captions and masks into spatial and relational dense captions. To accelerate the labeling process and improve caption quality, we present two VLM models: the Detailed Region Caption model and the Spatial Caption Merging model. Extensive experiments on various settings, including vision-language understanding, visual grounding, and region caption generation, demonstrate the effectiveness of our DenseWorld-1M dataset and labeling models.

多模态数据集视觉定位密集描述

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。