用频率感知规划解决图像修复多退化难题
FAPE-IR: Frequency-Aware Planning and Execution Framework for All-in-One Image Restoration
- 用冻结的多模态大模型做语义规划,生成频率导向修复指令
- 动态选择高低频专家模块,结合输入图像频率特征进行修复
- 适合需要统一处理多种图像退化的实际场景
全功能图像修复(AIO-IR)旨在构建一个能应对复杂条件下多种退化的统一模型。现有方法常依赖特定任务设计或隐式路由策略,难以适应真实世界中的多样化退化。本文提出频率感知规划与执行框架FAPE-IR,利用冻结的多模态大语言模型(MLLM)作为规划器,分析退化图像并生成简洁、频率敏感的修复计划。这些计划指导基于LoRA的专家混合(LoRA-MoE)模块在扩散模型执行器中动态选择高/低频专家,并融合输入图像的频率特征。为进一步提升修复质量并减少伪影,引入对抗训练和频率正则化损失。通过语义规划与频率驱动修复的耦合,FAPE-IR实现了统一且可解释的全功能图像修复。大量实验表明,该方法在七个修复任务上达到领先性能,并在混合退化下展现出强零样本泛化能力。
原文摘要 · Abstract (English)
All-in-One Image Restoration (AIO-IR) aims to develop a unified model that can handle multiple degradations under complex conditions. However, existing methods often rely on task-specific designs or latent routing strategies, making it hard to adapt to real-world scenarios with various degradations. We propose FAPE-IR, a Frequency-Aware Planning and Execution framework for image restoration. It uses a frozen Multimodal Large Language Model (MLLM) as a planner to analyze degraded images and generate concise, frequency-aware restoration plans. These plans guide a LoRA-based Mixture-of-Experts (LoRA-MoE) module within a diffusion-based executor, which dynamically selects high- or low-frequency experts, complemented by frequency features of the input image. To further improve restoration quality and reduce artifacts, we introduce adversarial training and a frequency regularization loss. By coupling semantic planning with frequency-based restoration, FAPE-IR offers a unified and interpretable solution for all-in-one image restoration. Extensive experiments show that FAPE-IR achieves state-of-the-art performance across seven restoration tasks and exhibits strong zero-shot generalization under mixed degradations.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。