用多模态大模型引导图像修复,提升复杂退化场景下的恢复效果。
Leveraging Multimodal Large Language Models for All-in-One Image Restoration via a Mixture of Frequency Experts

- 通过多模态大模型提取特征,指导低层修复过程。
- 引入频率专家混合模块,自适应组合不同频段修复专家。
- 在CDD11数据集上提升1.35 dB,适合复杂退化图像修复任务。
全功能图像修复旨在通过统一框架从受多种未知退化影响的输入中恢复清晰图像。现有方法通过识别退化特征来引导修复,但多将退化视为离散类别,难以建模复合退化中的连续关系结构。为此,我们提出一种基于多模态大语言模型(MLLM)的图像修复框架,利用多模态嵌入作为低层修复的引导信号。具体地,通过MLLM引导融合块(MGFB)将MLLM特征注入编码器-解码器架构,增强退化感知表示。同时,引入频率专家混合(MoFE)模块,利用MLLM引导的上下文线索自适应组合频率专家。为进一步优化专家路由,设计了具有关系对齐损失的MLLM引导路由机制,使路由模式与退化输入在嵌入空间中的关系保持一致。在多个基准上的大量实验表明,该方法在多样化的修复设置中表现优异,在挑战性数据集CDD11上达到新最优性能,相比先前方法最高提升1.35 dB。
原文摘要 · Abstract (English)
All-in-one image restoration seeks to recover clean images from inputs affected by diverse and unknown degradations using a unified framework. Recent methods have shown strong performance by identifying degradation characteristics to guide the restoration process. However, many of them treat degradations as discrete categories, which limits their ability to model the continuous relational structure that arises in composite degradations. To address this issue, we propose a multimodal large language model (MLLM)-guided image restoration framework that exploits multimodal embeddings as guidance for low-level restoration. Specifically, MLLM-derived features are injected into an encoder-decoder architecture through an MLLM-guided fusion block (MGFB) to enhance degradation-aware representations. In addition, we incorporate a mixture-of-frequency-experts (MoFE) module that adaptively combines frequency experts using MLLM-guided contextual cues. To further improve expert routing, we design an MLLM-guided router with a relational alignment loss that encourages routing patterns consistent with the embedding-space relationships of degraded inputs. Extensive experiments on multiple benchmarks show that the proposed method achieves strong performance across diverse restoration settings and establishes a new state of the art on the challenging CDD11 dataset, outperforming previous methods by up to 1.35 dB.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。