用自动数据生成让视觉语言模型精准理解复杂指令中的像素级指代。
Ground-V: Teaching VLMs to Ground Complex Instructions in Pixels
- 通过知识蒸馏自动生成带像素标注的指令-响应对,减少人工标注成本。
- 在6个基准上使LISA和PSALM的gIoU平均提升4.4%和7.9%。
- 特别适合需要细粒度视觉定位能力的研究者与开发者。
本文提出一种简单高效的自动化流程,用于扩展基于文本指令的指代消解数据,以激发视觉语言模型在复杂指令下的像素级定位能力。针对现实场景中五个关键挑战——幻觉引用、多对象、推理、多粒度及局部引用——我们利用预训练教师模型的知识蒸馏,生成高质量的指令-响应对,并关联已有像素级标注,显著降低人工标注成本。由此构建的数据集Ground-V包含丰富的物体定位知识与精细的像素级指代表达。实验表明,基于Ground-V训练的模型在多个定位任务中表现显著提升:在六项基准上,LISA与PSALM的gIoU平均分别提高4.4%和7.9%;在标准基准RefCOCO/+/g上达到新最优性能,其中gRefCOCO的N-Acc达83.3%,超越此前最佳结果超过20个百分点。
原文摘要 · Abstract (English)
This work presents a simple yet effective workflow for automatically scaling instruction-following data to elicit pixel-level grounding capabilities of VLMs under complex instructions. In particular, we address five critical real-world challenges in text-instruction-based grounding: hallucinated references, multi-object scenarios, reasoning, multi-granularity, and part-level references. By leveraging knowledge distillation from a pre-trained teacher model, our approach generates high-quality instruction-response pairs linked to existing pixel-level annotations, minimizing the need for costly human annotation. The resulting dataset, Ground-V, captures rich object localization knowledge and nuanced pixel-level referring expressions. Experiment results show that models trained on Ground-V exhibit substantial improvements across diverse grounding tasks. Specifically, incorporating Ground-V during training directly achieves an average accuracy boost of 4.4% for LISA and a 7.9% for PSALM across six benchmarks on the gIoU metric. It also sets new state-of-the-art results on standard benchmarks such as RefCOCO/+/g. Notably, on gRefCOCO, we achieve an N-Acc of 83.3%, exceeding the previous state-of-the-art by more than 20%.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。