arXiv:2411.13794cs.CV2024-11被引 3

构建大规模图像编辑数据集,提升生成模型在增删对象任务上的性能。

GalaxyEdit: Large-Scale Image Editing Dataset with Enhanced Diffusion Adapter

  • 自动化生成图像编辑数据,解决人工标注成本高问题。
  • 新模型在增删任务上FID分别提升11.2%和26.1%。
  • 轻量级适配器改进控制网络与U-Net通信,适合设备端使用。

大规模文本到图像及图像到图像模型的训练需要大量标注数据。尽管文本到图像数据集丰富,但基于指令的图像编辑数据(如物体添加与删除)仍稀缺,原因包括人力投入大、自动化程度低、端到端模型表现不佳、数据多样性受限及成本高昂。本文提出一种自动化数据生成流程,构建了大规模图像编辑数据集GalaxyEdit,支持添加与删除操作。在该数据集上微调SD v1.5模型后,新模型在处理更广泛物体和复杂指令时表现更优,添加与删除任务的FID分别降低11.2%和26.1%,优于现有方法。此外,为支持设备端应用,引入基于ControlNet-xs的轻量级适配器,并通过基于Volterra滤波器的非线性交互层增强控制网络与U-Net间的通信,显著提升复杂编辑任务性能,同时在边缘引导生成中也优于ControlNet-xs。

原文摘要 · Abstract (English)

Training of large-scale text-to-image and image-to-image models requires a huge amount of annotated data. While text-to-image datasets are abundant, data available for instruction-based image-to-image tasks like object addition and removal is limited. This is because of the several challenges associated with the data generation process, such as, significant human effort, limited automation, suboptimal end-to-end models, data diversity constraints and high expenses. We propose an automated data generation pipeline aimed at alleviating such limitations, and introduce GalaxyEdit - a large-scale image editing dataset for add and remove operations. We fine-tune the SD v1.5 model on our dataset and find that our model can successfully handle a broader range of objects and complex editing instructions, outperforming state-of-the-art methods in FID scores by 11.2\% and 26.1\% for add and remove tasks respectively. Furthermore, in light of on-device usage scenarios, we expand our research to include task-specific lightweight adapters leveraging the ControlNet-xs architecture. While ControlNet-xs excels in canny and depth guided generation, we propose to improve the communication between the control network and U-Net for more intricate add and remove tasks. We achieve this by enhancing ControlNet-xs with non-linear interaction layers based on Volterra filters. Our approach outperforms ControlNet-xs in both add/remove and canny-guided image generation tasks, highlighting the effectiveness of the proposed enhancement.

图像编辑扩散模型数据集轻量化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。