用少量样本和提示词,让大模型学会图像修复。
Edit2Restore:Few-Shot Image Restoration via Parameter-Efficient Adaptation of Pre-trained Editing Models
- 用轻量适配器微调预训练编辑模型,实现少样本图像修复。
- 仅需32-128张配对图像,效果超越百万级数据训练的基线。
- 统一适配器可处理五种不同退化类型,适合快速定制修复任务。
图像修复传统上需要为每种退化类型训练专用模型,依赖数千张配对数据。大型预训练文本条件图像编辑模型蕴含丰富的图像结构、质量与退化先验,但我们发现这些知识本身不足以直接用于修复:当前最先进编辑模型在零样本情况下表现不佳。我们证明其并非缺乏能力,而是缺乏方向;只需少量参数高效适配即可补足。在12B参数的FLUX.1 Kontext(一种图像到图像翻译的流匹配模型)上,仅用每任务32至128张配对图像,并通过简单文本提示引导,即可将原本表现平庸的零样本编辑器转化为具备竞争力的修复器。单一统一适配器,通过任务特定提示控制,可处理五种不同退化类型。尽管数据量仅为基线的三至四数量级,我们的少样本模型在多数感知与分布级指标上超越了基于超过一百万张精选配对数据训练的近期修复基线。评估聚焦于感知质量而非像素精度。通过全面分析训练集规模、任务特异性与统一多任务适配器的权衡、文本编码器适配的影响以及零样本基线性能,确立预训练编辑模型作为少样本、提示引导图像修复的数据高效基础。
原文摘要 · Abstract (English)
Image restoration has traditionally required training specialized models on thousands of paired examples per degradation type. Large pre-trained text-conditioned image editing models encode rich priors about image structure, quality, and degradation, yet we find that this knowledge does not, on its own, make them restorers: state-of-the-art editing models largely fail at restoration in the zero-shot regime. We show that what these priors lack is not capability but direction, and that a small amount of parameter-efficient adaptation supplies it. Fine-tuning LoRA adapters on FLUX.1 Kontext, a 12B-parameter flow matching model for image-to-image translation, with only 32--128 paired images per task and guided by simple text prompts, we turn a mediocre zero-shot editor into a competitive restorer. A single unified adapter, conditioned on task-specific prompts, handles five diverse degradations. Despite using three to four orders of magnitude less data, our few-shot model surpasses a recent restoration baseline trained on over a million curated pairs on the majority of perceptual and distribution-level metrics, on which we evaluate in keeping with our focus on perceptual rather than pixel-fidelity quality. Through comprehensive studies, we analyze the impact of training-set size, the trade-off between task-specific and unified multi-task adapters, the effect of text encoder adaptation, and zero-shot baseline performance, establishing pre-trained editing models as a compelling, data-efficient foundation for few-shot, prompt-guided image restoration.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。