用大模型生成数据,低成本实现高质量图像修复
Acquire and then Adapt: Squeezing out Text-to-Image Model for Image Restoration
- 用Flux模型自动生成海量修复训练数据
- 仅需8.5%成本达到顶尖修复效果
- 轻量适配器适合资源有限场景
近期预训练文本到图像(T2I)模型因强大的生成先验,被广泛用于真实世界图像修复。然而,控制这类大型模型进行修复通常需要大量高质量图像和巨大算力,成本高昂且不隐私友好。本文发现,已训练好的大型T2I模型(如Flux)能生成符合真实分布的多样化高质量图像,可作为无限训练样本来源。为此,我们提出FluxGen数据构建流程,包括无条件图像生成、图像筛选与退化模拟。同时设计轻量级适配器FluxIR,采用挤压-激励模块,有效控制基于Diffusion Transformer(DiT)的T2I模型以恢复合理细节。实验表明,该方法使Flux模型高效适配真实图像修复任务,在合成与真实退化数据集上均取得优异评分与视觉质量,训练成本仅为当前方法的约8.5%。
原文摘要 · Abstract (English)
Recently, pre-trained text-to-image (T2I) models have been extensively adopted for real-world image restoration because of their powerful generative prior. However, controlling these large models for image restoration usually requires a large number of high-quality images and immense computational resources for training, which is costly and not privacy-friendly. In this paper, we find that the well-trained large T2I model (i.e., Flux) is able to produce a variety of high-quality images aligned with real-world distributions, offering an unlimited supply of training samples to mitigate the above issue. Specifically, we proposed a training data construction pipeline for image restoration, namely FluxGen, which includes unconditional image generation, image selection, and degraded image simulation. A novel light-weighted adapter (FluxIR) with squeeze-and-excitation layers is also carefully designed to control the large Diffusion Transformer (DiT)-based T2I model so that reasonable details can be restored. Experiments demonstrate that our proposed method enables the Flux model to adapt effectively to real-world image restoration tasks, achieving superior scores and visual quality on both synthetic and real-world degradation datasets - at only about 8.5\% of the training cost compared to current approaches.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。