用视觉语言模型零训练生成带像素标签的合成图像,省去人工标注。
Data Factory with Minimal Human Effort Using VLMs
- 结合预训练ControlNet与VLM,无需训练即可生成带标签图像。
- 在PASCAL-5i和COCO-20i上实现优于现有方法的一次性语义分割性能。
- 适合需要快速构建高质量标注数据集的研究者使用。
通过数据增强生成足够且多样化的数据,可有效缓解收集和标注像素级图像所耗费的时间与人力。传统数据增强方法在操控高阶语义属性(如材质、纹理)方面存在困难。相比之下,扩散模型可通过文本到图像或图像到图像转换提供稳健替代方案。然而,现有基于扩散的方法要么计算成本高,要么性能妥协。为此,我们提出一种无需训练的新流程,整合预训练ControlNet与视觉语言模型(VLMs),生成配对的合成图像及像素级标签。该方法无需手动标注,显著提升下游任务表现。为提高图像保真度与多样性,我们引入多路提示生成器、掩码生成器与高质量图像选择模块。在PASCAL-5i与COCO-20i上的实验表明,该方法在一次性语义分割任务中表现优异,超越同期工作。
原文摘要 · Abstract (English)
Generating enough and diverse data through augmentation offers an efficient solution to the time-consuming and labour-intensive process of collecting and annotating pixel-wise images. Traditional data augmentation techniques often face challenges in manipulating high-level semantic attributes, such as materials and textures. In contrast, diffusion models offer a robust alternative, by effectively utilizing text-to-image or image-to-image transformation. However, existing diffusion-based methods are either computationally expensive or compromise on performance. To address this issue, we introduce a novel training-free pipeline that integrates pretrained ControlNet and Vision-Language Models (VLMs) to generate synthetic images paired with pixel-level labels. This approach eliminates the need for manual annotations and significantly improves downstream tasks. To improve the fidelity and diversity, we add a Multi-way Prompt Generator, Mask Generator and High-quality Image Selection module. Our results on PASCAL-5i and COCO-20i present promising performance and outperform concurrent work for one-shot semantic segmentation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。