优化图像修复模型的训练数据配比,提升多任务泛化能力。
FoundIR-v2: Optimizing Pre-Training Data Mixtures for Image Restoration Foundation Model
- 动态调整不同修复任务的数据混合比例,实现训练平衡。
- 支持超过50个子任务,在真实场景中表现优于现有方法。
- 采用专家混合调度器,自适应分配不同退化类型的先验知识。
近期研究显示,图像修复基础模型的性能提升主要依赖于预训练数据规模与质量的增强。本文发现,不同修复任务间的数据混合比例同样是决定全任务统一模型整体性能的关键因素。为此,我们提出高容量的基于扩散的图像修复基础模型 FoundIR-v2,采用数据均衡调度策略,动态优化来自不同任务的训练数据混合比例。通过利用数据混合规律,该方法确保数据集构成均衡,使模型在多样任务上实现一致的泛化能力和全面性能。此外,我们在生成式预训练中引入有效的专家混合(MoE)驱动调度器,灵活为每类修复任务分配适配的扩散先验,以应对不同任务间的退化形式与程度差异。大量实验表明,本方法可覆盖更广泛的真实世界场景中的50多个子任务,并在性能上优于当前最优方法。
原文摘要 · Abstract (English)
Recent studies have witnessed significant advances in image restoration foundation models driven by improvements in the scale and quality of pre-training data. In this work, we find that the data mixture proportions from different restoration tasks are also a critical factor directly determining the overall performance of all-in-one image restoration models. To this end, we propose a high-capacity diffusion-based image restoration foundation model, FoundIR-v2, which adopts a data equilibrium scheduling paradigm to dynamically optimize the proportions of mixed training datasets from different tasks. By leveraging the data mixing law, our method ensures a balanced dataset composition, enabling the model to achieve consistent generalization and comprehensive performance across diverse tasks. Furthermore, we introduce an effective Mixture-of-Experts (MoE)-driven scheduler into generative pre-training to flexibly allocate task-adaptive diffusion priors for each restoration task, accounting for the distinct degradation forms and levels exhibited by different tasks. Extensive experiments demonstrate that our method can address over 50 sub-tasks across a broader scope of real-world scenarios and achieves favorable performance against state-of-the-art approaches.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。