用强化学习让图像修复更高效,无需标注也能自动选修复步骤。
Restore-R1: Efficient Image Restoration Agents via Reinforcement Learning with Multimodal LLM Perceptual Feedback
- 通过强化学习训练轻量代理,按步选择最优修复操作
- 在无标注环境下达到顶尖修复效果,且推理速度显著提升
- 适合需要快速修复且无标签数据的场景
复杂图像修复旨在从受模糊、噪声、雨痕和压缩伪影等多种退化影响的输入中恢复高质量图像。近期基于视觉-语言模型和大语言模型的修复代理虽具潜力,但因反思、回滚和迭代工具搜索导致效率低下。此外,其性能高度依赖需大量标注训练的退化识别模型,在无标签环境中应用受限。为此,我们提出一种基于策略优化的修复框架,训练一个轻量级代理以确定工具调用序列。该代理在顺序决策过程中每步选择最合适的修复操作,以最大化最终图像质量。为实现无标注环境下的训练,我们引入由多模态大语言模型驱动的新奖励机制,充当人类对齐评估者,提供感知反馈以优化策略。训练完成后,代理执行确定性修复计划,避免冗余调用,显著加速推理同时保持高修复质量。大量实验表明,尽管无监督训练,本方法在全参考指标上达到当前最佳水平,并在多种退化场景下超越现有方法的无参考指标表现。
原文摘要 · Abstract (English)
Complex image restoration aims to recover high-quality images from inputs affected by multiple degradations such as blur, noise, rain, and compression artifacts. Recent restoration agents, powered by vision-language models and large language models, offer promising restoration capabilities but suffer from significant efficiency bottlenecks due to reflection, rollback, and iterative tool searching. Moreover, their performance heavily depends on degradation recognition models that require extensive annotations for training, limiting their applicability in label-free environments. To address these limitations, we propose a policy optimization-based restoration framework that learns an lightweight agent to determine tool-calling sequences. The agent operates in a sequential decision process, selecting the most appropriate restoration operation at each step to maximize final image quality. To enable training within label-free environments, we introduce a novel reward mechanism driven by multimodal large language models, which act as human-aligned evaluator and provide perceptual feedback for policy improvement. Once trained, our agent executes a deterministic restoration plans without redundant tool invocations, significantly accelerating inference while maintaining high restoration quality. Extensive experiments show that despite using no supervision, our method matches SOTA performance on full-reference metrics and surpasses existing approaches on no-reference metrics across diverse degradation scenarios.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。