自动生成图文指代理解与生成数据,省去人工标注。
ColLab: A Collaborative Spatial Progressive Data Engine for Referring Expression Comprehension and Generation
- 用多模态大模型协作生成语言描述,无需人工干预。
- 通过空间渐进增强,提升重复物体的表达区分度。
- 已用于ICCV 2025挑战赛数据生成,适配真实场景需求。
指代表达理解(REC)与生成(REG)是多模态理解的核心任务,可通过自然语言实现精确目标定位。然而,现有REC与REG数据集严重依赖人工标注,成本高且难以扩展。本文提出ColLab,一种无需人工监督的协同空间渐进式数据生成引擎,可全自动完成REC与REG数据生成。方法引入协同多模态模型交互(CMMI)策略,利用多模态大语言模型(MLLM)和大语言模型(LLM)的语义理解能力生成描述;同时设计空间渐进增强(SPA)模块,提升重复实例间的空间表达能力。实验表明,ColLab显著加速标注流程,同时提升生成表达的质量与区分度。此外,本框架部分被采纳至ICCV 2025 MARS2挑战赛的数据生成流程中,丰富了多样且具挑战性的样本,更贴近真实世界推理需求。
原文摘要 · Abstract (English)
Referring Expression Comprehension (REC) and Referring Expression Generation (REG) are fundamental tasks in multimodal understanding, supporting precise object localization through natural language. However, existing REC and REG datasets rely heavily on manual annotation, which is labor-intensive and difficult to scale. In this paper, we propose ColLab, a collaborative spatial progressive data engine that enables fully automated REC and REG data generation without human supervision. Specifically, our method introduces a Collaborative Multimodal Model Interaction (CMMI) strategy, which leverages the semantic understanding of multimodal large language models (MLLMs) and large language models (LLMs) to generate descriptions. Furthermore, we design a module termed Spatial Progressive Augmentation (SPA) to enhance spatial expressiveness among duplicate instances. Experiments demonstrate that ColLab significantly accelerates the annotation process of REC and REG while improving the quality and discriminability of the generated expressions. In addition to the core methodological contribution, our framework was partially adopted in the data generation pipeline of the ICCV 2025 MARS2 Challenge on Multimodal Reasoning, enriching the dataset with diverse and challenging samples that better reflect real-world reasoning demands.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。