通过智能重述编辑任务,让图像修改更可靠。
Making Image Editing Easier via Adaptive Task Reformulation with Agentic Executions

- 用多模态大模型动态生成编辑步骤,自动优化指令
- 在多个数据集上提升效果,尤其对复杂任务增益显著
- 适合希望提升图像编辑成功率的研究者和开发者
尽管基于生成模型的指令引导图像编辑已取得显著进展,但在许多看似简单的场景中仍难以产出可靠结果。我们发现,这些失败主要源于任务描述不佳,如目标过小、空间关系隐含或指令不明确。本文将图像编辑失败视为任务表述问题,提出一种无需修改底层模型的自适应任务重述框架。核心思想是利用多模态大模型(MLLM)代理,通过分析、路由、重述与反馈驱动的迭代优化,将原始图像-指令对转化为一系列动态生成的操作序列。在ImgEdit、PICA和RePlan等多个基准上,针对Qwen Image Edit和Nano Banana等多种编辑模型的实验均显示一致性能提升,尤其在挑战性任务上增益明显。结果表明,任务重述是关键但被忽视的影响因素,通过更好匹配模型有效工作区,可实现显著性能提升。
原文摘要 · Abstract (English)
Instruction guided image editing has advanced substantially with recent generative models, yet it still fails to produce reliable results across many seemingly simple cases. We observe that a large portion of these failures stem not from insufficient model capacity, but from poorly formulated editing tasks, such as those involving small targets, implicit spatial relations, or under-specified instructions. In this work, we frame image editing failures as a task formulation problem and propose an adaptive task reformulation framework that improves editing performance without modifying the underlying model. Our key idea is to transform the original image-instruction pair into a sequence of operations that are dynamically determined and executed by a MLLM agent through analysis, routing, reformulation, and feedback-driven refinement. Experiments on multiple benchmarks, including ImgEdit, PICA, and RePlan, across diverse editing backbones such as Qwen Image Edit and Nano Banana, show consistent improvements, with especially large gains on challenging cases. These results suggest that task reformulation is a critical but underexplored factor, and that substantial gains can be achieved by better matching editing tasks to the effective operating regime of existing models.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。