让图像修复智能体通过经验自我进化,减少试错,提升效果与效率。
EvoIR-Agent: Self-Evolving Image Restoration Agentic System via Experience-Driven Learning

- 构建分层经验池,指导工具选择与去噪顺序。
- 自进化机制从零积累记录,性能超越当前最优方法。
- 适合需要快速适配新工具或退化类型的图像修复场景。
基于多模态大语言模型的图像修复智能体在退化耦合场景中表现优异,能灵活选择工具并确定去除顺序。然而,零样本规划常因缺乏经验而失败,导致大量试错开销。现有方法面临两难:训练型方法将经验嵌入参数,推理高效但难以兼容新工具或退化类型;无训练方法虽可显式存储经验,仍因经验粗糙导致试错成本高。为此,我们提出EvoIR-Agent,首次系统性定义无训练图像修复智能体的经验构成。构建分层经验池,实现对多种工具和去除顺序的粗粒度到细粒度引导。引入自进化机制,利用累积记录从零更新经验池,显著提升性能与效率。大量实验表明,EvoIR-Agent在全参考指标上取得显著领先,并在性能与效率间实现卓越的帕累托最优平衡,优于当前最先进方法。
原文摘要 · Abstract (English)
Multimodal Large Language Model (MLLM)-driven image restoration agent demonstrates effectiveness in degradation coupling scenarios by flexibly selecting tools and determining removal orders. However, their zero-shot planning often fails without experience, necessitating severe trial-and-error overhead to achieve satisfactory outcomes. Currently, two paradigms are employed to address this issue, yet a dilemma persists: Training-based methods embed intrinsic experience into parameters, achieving high inference efficiency but lacking compatibility with new tools or degradation. In contrast, training-free methods utilize explicit experience storage for compatibility but still incur trial-and-error overhead due to naive experience. To resolve the dilemma, we propose EvoIR-Agent, which first systematically formulates the experience components of a training-free image restoration agent. Subsequently, a hierarchical experience pool is constructed, which enables coarse-to-fine guidance for diverse tools and removal orders. Furthermore, a self-evolving mechanism is introduced to update the pool from scratch using accumulated records, thereby greatly improving performance and efficiency. Extensive experiments reveal that EvoIR-Agent achieves a significant lead in the full reference metrics and yields a remarkable Pareto-optimal balance between performance and efficiency compared to the state-of-the-art methods.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。