用大模型拆解复杂指令,自动规划精准图像编辑
Beyond Simple Edits: X-Planner for Complex Instruction-Based Image Editing
- 用思维链分解复杂指令为可执行子任务
- 自动生成精确编辑类型与掩码,保持人物身份一致
- 无需手动画框,适合复杂指令编辑场景
基于扩散模型的图像编辑方法虽在文本引导任务上取得进展,但常难以理解复杂、间接的指令,且存在身份丢失、误编辑或依赖人工掩码的问题。为此,我们提出 X-Planner,一种基于多模态大语言模型的规划系统,能将用户意图与编辑能力有效对齐。X-Planner通过思维链推理,系统性地将复杂指令分解为清晰的子指令,并为每个子指令自动生成准确的编辑类型和分割掩码,避免人工干预,实现局部化、身份保留的编辑。此外,我们设计了一种新型自动化数据生成流程,用于训练 X-Planner,使其在现有基准和新提出的复杂编辑基准上均达到领先性能。
原文摘要 · Abstract (English)
Recent diffusion-based image editing methods have significantly advanced text-guided tasks but often struggle to interpret complex, indirect instructions. Moreover, current models frequently suffer from poor identity preservation, unintended edits, or rely heavily on manual masks. To address these challenges, we introduce X-Planner, a Multimodal Large Language Model (MLLM)-based planning system that effectively bridges user intent with editing model capabilities. X-Planner employs chain-of-thought reasoning to systematically decompose complex instructions into simpler, clear sub-instructions. For each sub-instruction, X-Planner automatically generates precise edit types and segmentation masks, eliminating manual intervention and ensuring localized, identity-preserving edits. Additionally, we propose a novel automated pipeline for generating large-scale data to train X-Planner which achieves state-of-the-art results on both existing benchmarks and our newly introduced complex editing benchmark.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。