用强化学习协调多个专业智能体,实现更精准的多步图像编辑。
ImageEdit-R1: Boosting Multi-Agent Image Editing via Reinforcement Learning
- 将图像编辑拆解为多智能体协作的决策过程,各司其职。
- 在多个数据集上超越闭源扩散模型和现有多智能体框架。
- 适合需要复杂、多步骤指令的图像编辑场景。
随着商用多模态模型的快速发展,图像编辑因其在日常生活中的广泛应用而备受关注。尽管取得了显著进展,现有的图像编辑系统(尤其是闭源或专有模型)在处理复杂、间接或多步骤用户指令时仍存在困难,难以实现与人类意图一致的精细、上下文感知编辑。本文提出 ImageEdit-R1,一种基于强化学习的多智能体图像编辑框架,通过协调一组具备特定能力的预训练视觉-语言与生成智能体来实现高阶决策。每个智能体负责理解用户意图、定位感兴趣区域、选择合适编辑动作及合成视觉内容,强化学习则驱动它们协同工作,确保行为连贯且目标导向。与依赖单体模型或手工设计流水线的方法不同,本方法将图像编辑视为序列决策问题,支持动态、上下文感知的编辑策略。实验结果表明,ImageEdit-R1 在多个图像编辑数据集上持续优于单一闭源扩散模型及其它多智能体框架基线。
原文摘要 · Abstract (English)
With the rapid advancement of commercial multi-modal models, image editing has garnered significant attention due to its widespread applicability in daily life. Despite impressive progress, existing image editing systems, particularly closed-source or proprietary models, often struggle with complex, indirect, or multi-step user instructions. These limitations hinder their ability to perform nuanced, context-aware edits that align with human intent. In this work, we propose ImageEdit-R1, a multi-agent framework for intelligent image editing that leverages reinforcement learning to coordinate high-level decision-making across a set of specialized, pretrained vision-language and generative agents. Each agent is responsible for distinct capabilities--such as understanding user intent, identifying regions of interest, selecting appropriate editing actions, and synthesizing visual content--while reinforcement learning governs their collaboration to ensure coherent and goal-directed behavior. Unlike existing approaches that rely on monolithic models or hand-crafted pipelines, our method treats image editing as a sequential decision-making problem, enabling dynamic and context-aware editing strategies. Experimental results demonstrate that ImageEdit-R1 consistently outperforms both individual closed-source diffusion models and alternative multi-agent framework baselines across multiple image editing datasets.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。