用统一框架生成可控图像数据,提升视觉语言任务泛化能力
AnySynth: Harnessing the Power of Image Synthetic Data Generation for Generalized Vision-Language Tasks
- 基于大模型和布局先验生成任务适配的图像布局
- 可生成高质量、可控制的合成图像,支持多任务需求
- 适合需要大量标注数据的视觉语言研究者使用
扩散模型近期被用于生成高质量图像,减少人工数据收集需求,并提升目标检测、实例分割和图像感知等任务的模型泛化能力。然而,现有合成框架通常需针对每项任务精心设计,受限于图像布局、内容与标注格式差异,难以推广到通用场景。本文提出 AnySynth,一个统一框架,集成可适配、全面且高度可控的组件,能根据多样需求生成任意类型的合成数据。首先引入任务特定布局生成模块,利用大语言模型生成能力和真实图像布局先验,生成合理布局;随后开发统一可控图像生成模块,基于生成布局创建高质量合成图像,并支持用户提供的参考图与风格图;最后通过任务导向标注模块为不同任务提供精确详尽的标注。我们在少样本目标检测、跨域目标检测、零样本组合图像检索及多模态图像感知与定位等多个任务上验证了该框架性能,合成数据显著提升了模型表现,证明其通用性与有效性。
原文摘要 · Abstract (English)
Diffusion models have recently been employed to generate high-quality images, reducing the need for manual data collection and improving model generalization in tasks such as object detection, instance segmentation, and image perception. However, the synthetic framework is usually designed with meticulous human effort for each task due to various requirements on image layout, content, and annotation formats, restricting the application of synthetic data on more general scenarios. In this paper, we propose AnySynth, a unified framework integrating adaptable, comprehensive, and highly controllable components capable of generating an arbitrary type of synthetic data given diverse requirements. Specifically, the Task-Specific Layout Generation Module is first introduced to produce reasonable layouts for different tasks by leveraging the generation ability of large language models and layout priors of real-world images. A Uni-Controlled Image Generation Module is then developed to create high-quality synthetic images that are controllable and based on the generated layouts. In addition, user specific reference images, and style images can be incorporated into the generation to task requirements. Finally, the Task-Oriented Annotation Module offers precise and detailed annotations for the generated images across different tasks. We have validated our framework's performance across various tasks, including Few-shot Object Detection, Cross-domain Object Detection, Zero-shot Composed Image Retrieval, and Multi-modal Image Perception and Grounding. The specific data synthesized by our framework significantly improves model performance in these tasks, demonstrating the generality and effectiveness of our framework.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。