arXiv:2411.14807cs.CVcs.CL2024-11中稿 · ICPR 2024被引 3

用颜色驱动生成百万级合成数据,提升视觉语言理解性能。

Harlequin: Color-driven Generation of Synthetic Data for Referring Expression Comprehension

  • 基于颜色与语义变化生成多样化的文本-图像配对数据
  • 构建超百万条查询的合成数据集,显著降低标注成本
  • 适合需要大规模训练数据的视觉语言模型研究者

指代表达理解(REC)旨在通过自然语言描述定位场景中的特定物体,是视觉语言理解的重要任务。当前主流方法依赖深度学习,通常需昂贵且人工标注的数据。部分工作尝试在弱监督或借助大视觉语言模型的前提下解决此问题,但合成标注数据的技术发展被忽视。本文提出一种新框架,可同时考虑文本与视觉模态,生成适用于REC任务的人工数据。首先,该流水线对现有数据进行注释变异处理;随后,以修改后的注释为引导生成新图像。最终产出一个包含超过100万条查询的新数据集,名为Harlequin。该方法免去人工采集与标注,具备高可扩展性,并支持任意复杂度。我们在Harlequin上预训练三个REC模型,再在人工标注数据集上微调与评估。实验表明,基于人工数据的预训练能有效提升模型性能。

原文摘要 · Abstract (English)

Referring Expression Comprehension (REC) aims to identify a particular object in a scene by a natural language expression, and is an important topic in visual language understanding. State-of-the-art methods for this task are based on deep learning, which generally requires expensive and manually labeled annotations. Some works tackle the problem with limited-supervision learning or relying on Large Vision and Language Models. However, the development of techniques to synthesize labeled data is overlooked. In this paper, we propose a novel framework that generates artificial data for the REC task, taking into account both textual and visual modalities. At first, our pipeline processes existing data to create variations in the annotations. Then, it generates an image using altered annotations as guidance. The result of this pipeline is a new dataset, called Harlequin, made by more than 1M queries. This approach eliminates manual data collection and annotation, enabling scalability and facilitating arbitrary complexity. We pre-train three REC models on Harlequin, then fine-tuned and evaluated on human-annotated datasets. Our experiments show that the pre-training on artificial data is beneficial for performance.

视觉语言数据合成指代理解生成模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。