构建百万级真实图像修复数据集,推动通用图像修复模型发展
FoundIR: Unleashing Million-scale Training Data to Advance Foundation Models for Image Restoration
- 用多轮采集与对齐准则构建百万级真实图像对数据集
- 提出分阶段训练的FoundIR模型,在复杂场景下实现高质量修复
- 适合图像修复、基础模型研究者参考
尽管全能型模型在通用图像修复上取得进展,但现有方法在真实场景中仍面临泛化瓶颈,主要因训练数据规模小且退化类型单一。为此,我们贡献了一个百万级数据集,具有两大优势:更大规模的真实样本和更丰富的退化类型。通过调整相机参数与外部成像条件,利用设计的数据采集系统和对齐准则,多轮获取对齐图像对。同时,提出鲁棒模型FoundIR,进一步迈向图像修复基础模型。首先采用基于扩散的通用模型,通过学习多样化输入中的退化无关共性表征进行修复,并引入增量学习策略优化训练;为提升复杂场景下的修复能力,再引入退化感知专家模型以获得高质量结果。大量实验验证了数据集价值与方法有效性。
原文摘要 · Abstract (English)
Despite the significant progress made by all-in-one models in universal image restoration, existing methods suffer from a generalization bottleneck in real-world scenarios, as they are mostly trained on small-scale synthetic datasets with limited degradations. Therefore, large-scale high-quality real-world training data is urgently needed to facilitate the emergence of foundational models for image restoration. To advance this field, we spare no effort in contributing a million-scale dataset with two notable advantages over existing training data: real-world samples with larger-scale, and degradation types with higher diversity. By adjusting internal camera settings and external imaging conditions, we can capture aligned image pairs using our well-designed data acquisition system over multiple rounds and our data alignment criterion. Moreover, we propose a robust model, FoundIR, to better address a broader range of restoration tasks in real-world scenarios, taking a further step toward foundation models. Specifically, we first utilize a diffusion-based generalist model to remove degradations by learning the degradation-agnostic common representations from diverse inputs, where incremental learning strategy is adopted to better guide model training. To refine the model's restoration capability in complex scenarios, we introduce degradation-aware specialist models for achieving final high-quality results. Extensive experiments show the value of our dataset and the effectiveness of our method.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。