让图像修复模型像医生一样诊断问题再修复,提升效果与可解释性。
RIRF: Reasoning Image Restoration Framework

- 用大模型分析图像退化类型、严重程度和场景语义,生成诊断报告。
- 将诊断结果量化为强化学习信号,驱动修复模块精准优化。
- 统一框架实现推理与修复联动,比现有方法更准更透明。
通用图像修复(UIR)旨在使用统一模型从多种未知退化中恢复清晰图像。现有方法主要关注像素重建,缺乏对退化类型、严重程度及场景语义的显式诊断推理。我们提出“推理与修复”(R&R)框架,将结构化思维链(CoT)融入修复流程。R&R通过微调Qwen3-VL构建显式推理器,诊断退化类型、量化退化严重程度、推断关键退化因素,并描述相关场景与物体语义。生成的结构化推理为修复器提供可解释的细粒度先验。为进一步提升修复质量,推理器输出的退化严重度被用作强化学习(RL)信号,指导并增强修复器。不同于将推理与低层视觉任务分离的多模态大模型代理系统,R&R在统一框架中紧密耦合语义诊断与像素级修复。跨多个UIR基准的大量实验表明,R&R达到业界领先性能,同时提供对修复过程的独特可解释性。
原文摘要 · Abstract (English)
Universal image restoration (UIR) aims to recover clean images from diverse and unknown degradations using a unified model. Existing UIR methods primarily focus on pixel reconstruction and often lack explicit diagnostic reasoning over degradation composition, severity, and scene semantics prior to restoration. We propose Reason and Restore (R\&R), a novel framework that integrates structured Chain-of-Thought (CoT) reasoning into the image restoration pipeline. R\&R introduces an explicit reasoner, implemented by fine-tuning Qwen3-VL, to diagnose degradation types, quantify degradation severity, infer key degradation-related factors, and describe relevant scene and object semantics. The resulting structured reasoning provides interpretable and fine-grained diagnostic priors for the restorer. To further improve restoration quality, the quantified degradation severity produced by the reasoner is leveraged as reinforcement learning (RL) signals to guide and strengthen the restorer. Unlike existing multimodal LLM-based agentic systems that decouple reasoning from low-level vision tasks, R\&R tightly couples semantic diagnostic reasoning with pixel-level restoration in a unified framework. Extensive experiments across diverse UIR benchmarks demonstrate that R\&R achieves state-of-the-art performance while offering unique interpretability into the restoration process.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。