构建可追踪物体身份的图文生成数据集,让描述更可信
GroundCap: A Visually Grounded Image Captioning Dataset
- 用唯一ID追踪物体,实现动作与对象的精准关联
- 涵盖52,016张电影图,支持动作-物体联合标注
- 适合需要可验证视觉描述的研究者和开发者
当前图像描述系统难以将文本与具体视觉元素对应,导致输出难以验证。现有方法虽具备一定定位能力,但无法跨引用追踪物体身份,也无法同时定位动作与对象。我们提出一种基于ID的定位机制,实现物体身份持续追踪与动作-对象关联。构建GroundCap数据集,包含77部电影中的52,016张图像,提供344条人工标注与52,016条自动生成的带标注描述。每条描述通过标签系统关联检测到的132类物体与51类动作,并保持物体身份一致性;通过K-means聚类分割背景。提出gMETEOR评估指标,融合描述质量与定位精度;在该数据集上微调Pixtral-12B与Qwen2.5-VL 7B模型建立基线。人工评估表明,该方法能生成具有连贯物体指代、可验证的描述。
原文摘要 · Abstract (English)
Current image captioning systems lack the ability to link descriptive text to specific visual elements, making their outputs difficult to verify. While recent approaches offer some grounding capabilities, they cannot track object identities across multiple references or ground both actions and objects simultaneously. We propose a novel ID-based grounding system that enables consistent object reference tracking and action-object linking. We present GroundCap, a dataset containing 52,016 images from 77 movies, with 344 human-annotated and 52,016 automatically generated captions. Each caption is grounded on detected objects (132 classes) and actions (51 classes) using a tag system that maintains object identity while linking actions to the corresponding objects. Our approach features persistent object IDs for reference tracking, explicit action-object linking, and the segmentation of background elements through K-means clustering. We propose gMETEOR, a metric combining caption quality with grounding accuracy, and establish baseline performance by fine-tuning Pixtral-12B and Qwen2.5-VL 7B on GroundCap. Human evaluation demonstrates our approach's effectiveness in producing verifiable descriptions with coherent object references.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。