让图像编辑器像人一样反复思考,提升指令理解能力。
EditThinker: Unlocking Iterative Reasoning for Any Image Editor
- 引入思考-编辑循环,迭代优化指令与结果。
- 在四个基准上显著提升指令遵循率,效果明显优于现有方法。
- 适用于任何图像编辑模型,适合希望提升编辑精度的研究者。
基于指令的图像编辑已成为重要研究方向,得益于图像生成基础模型,已实现高美学质量,但指令遵循能力成为主要挑战。现有方法通过监督或强化学习提升指令对齐,然而单轮成功率受限于固有随机性与缺乏反思。本文提出一种反思式编辑框架,模拟人类认知过程,通过持续执行‘评估结果-优化指令’的循环,直至满意。我们训练一个统一的多模态大模型 EditThinker 作为推理引擎,联合生成评判分数、推理过程与优化指令。采用强化学习使模型思考与编辑行为对齐,从而生成更精准的指令改进。在四个基准上的大量实验表明,该方法显著提升任意图像编辑模型的指令遵循能力。我们将发布数据构建框架、数据集与模型以促进社区发展。
原文摘要 · Abstract (English)
Instruction-based image editing has emerged as a prominent research area, which, benefiting from image generation foundation models, have achieved high aesthetic quality, making instruction-following capability the primary challenge. Existing approaches improve instruction adherence via supervised or reinforcement learning, yet single-turn success rates remain limited due to inherent stochasticity and a lack of deliberation. In this work, we propose a deliberative editing framework to 'think' while they edit, which simulates the human cognitive loop by iteratively executing a Think-while-Edit cycle: Critiquing results and Refining instructions , followed by Repeating the generation until satisfactory. Specifically, we train a single MLLM, EditThinker, to act as the reasoning engine of this framework, which jointly produce the critique score, reasoning process, and refined instructions. We employ reinforcement learning to align the EditThinker's thinking with its editing, thereby generating more targeted instruction improvements. Extensive experiments on four benchmarks demonstrate that our approach significantly improves the instruction-following capability of any image editing model by a large margin. We will release our data construction framework, datasets, and models to benefit the community.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。