让AI理解复杂编辑指令,自动规划并执行多步图像修改。
From Plans to Pixels: Learning to Plan and Orchestrate for Open-Ended Image Editing

- 用计划器拆解抽象指令,调度器按步骤选工具和区域执行。
- 在12个复杂指令上达到87%的指令遵循率,优于基线模型。
- 适合需要长程推理与真实视觉质量的开放域图像编辑场景。
现代图像编辑模型虽能生成逼真结果,但在处理抽象、多步骤指令(如“让这张广告更符合素食主义”)时表现不佳。已有基于代理的方法虽可分解任务,但依赖手工设计流程或教师模仿,灵活性差且学习与实际编辑结果脱节。本文提出一种经验式长程图像编辑框架:计划器生成结构化的原子步骤分解,调度器选择工具与操作区域执行每一步;视觉语言判别器提供基于结果的奖励,衡量指令遵循度与视觉质量。调度器通过最大化奖励进行训练,成功轨迹用于优化计划器。通过将规划与基于奖励的执行紧密耦合,该方法在多个复杂指令上生成更连贯可靠的编辑结果,显著优于单步或规则驱动的多步基线模型。
原文摘要 · Abstract (English)
Modern image editing models produce realistic results but struggle with abstract, multi step instructions (e.g., ``make this advertisement more vegetarian-friendly''). Prior agent based methods decompose such tasks but rely on handcrafted pipelines or teacher imitation, limiting flexibility and decoupling learning from actual editing outcomes. We propose an experiential framework for long-horizon image editing, where a planner generates structured atomic decompositions and an orchestrator selects tools and regions to execute each step. A vision language judge provides outcome-based rewards for instruction adherence and visual quality. The orchestrator is trained to maximize these rewards, and successful trajectories are used to refine the planner. By tightly coupling planning with reward driven execution, our approach yields more coherent and reliable edits than single-step or rule-based multistep baselines.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。