arXiv:2502.02589cs.CV2025-02NeurIPS被引 13

构建细粒度图像描述数据集,提升视觉语言模型理解与生成能力

COCONut-PanCap: Joint Panoptic Segmentation and Grounded Captions for Fine-Grained Understanding and Generation

论文配图:COCONut-PanCap: Joint Panoptic Segmentation and Grounded Captions for Fine-Grained Understanding and Generation
图 1 · 摘自论文原文
  • 基于COCONut分割掩码,为每个区域生成精准图文描述
  • 人类精标密集标注,显著提升模型在理解与生成任务上的性能
  • 适合研究多模态理解、图像生成与细粒度描述的学者使用

本文提出COCONut-PanCap数据集,基于COCO并融合先进的COCONut全景分割掩码,旨在克服现有图像-文本数据集在细节和场景完整性上的不足。该数据集包含细粒度、区域级的图像描述,与全景分割掩码对齐,确保描述一致性与高细节性。通过人工编辑的密集标注,支持视觉语言模型(VLMs)在图像理解与文本到图像生成任务中的训练。实验表明,该数据集显著提升多项任务性能,为联合全景分割与接地描述任务提供新基准,推动多模态学习中高质量图像-文本标注的发展。

原文摘要 · Abstract (English)

This paper introduces the COCONut-PanCap dataset, created to enhance panoptic segmentation and grounded image captioning. Building upon the COCO dataset with advanced COCONut panoptic masks, this dataset aims to overcome limitations in existing image-text datasets that often lack detailed, scene-comprehensive descriptions. The COCONut-PanCap dataset incorporates fine-grained, region-level captions grounded in panoptic segmentation masks, ensuring consistency and improving the detail of generated captions. Through human-edited, densely annotated descriptions, COCONut-PanCap supports improved training of vision-language models (VLMs) for image understanding and generative models for text-to-image tasks. Experimental results demonstrate that COCONut-PanCap significantly boosts performance across understanding and generation tasks, offering complementary benefits to large-scale datasets. This dataset sets a new benchmark for evaluating models on joint panoptic segmentation and grounded captioning tasks, addressing the need for high-quality, detailed image-text annotations in multi-modal learning.

多模态图像描述分割数据集

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。