arXiv:2602.22624cs.CVcs.AI2026-02ICCV被引 6

用多模态思维链提升指令图像编辑的准确性和复杂场景适应性

Instruction-based Image Editing with Planning, Reasoning, and Generation

  • 分步构建思维链:规划、区域推理、编辑三阶段协同
  • 在真实复杂图像上实现媲美顶尖方法的编辑效果
  • 适合需要高精度语义理解的交互式图像编辑应用

通过指令进行图像编辑是一种自然的交互内容生成方式,但因对场景理解与生成能力要求更高而极具挑战。现有方法依赖大语言模型、目标分割模型和编辑模型组成的链条,但理解模块仅支持单一模态,限制了编辑质量。本文提出一种新型多模态模型,通过分离指令编辑任务为三个阶段:思维链(CoT)规划、编辑区域推理与图像编辑。在规划阶段,大语言模型根据指令和编辑网络能力生成合适子提示;在区域推理阶段,训练基于指令的多模态大语言模型以生成编辑区域;最后,设计一个提示引导的指令编辑网络,基于大规模文本到图像扩散模型实现高质量生成。大量实验表明,该方法在复杂真实图像上具备优异的编辑能力。

原文摘要 · Abstract (English)

Editing images via instruction provides a natural way to generate interactive content, but it is a big challenge due to the higher requirement of scene understanding and generation. Prior work utilizes a chain of large language models, object segmentation models, and editing models for this task. However, the understanding models provide only a single modality ability, restricting the editing quality. We aim to bridge understanding and generation via a new multi-modality model that provides the intelligent abilities to instruction-based image editing models for more complex cases. To achieve this goal, we individually separate the instruction editing task with the multi-modality chain of thought prompts, i.e., Chain-of-Thought (CoT) planning, editing region reasoning, and editing. For Chain-of-Thought planning, the large language model could reason the appropriate sub-prompts considering the instruction provided and the ability of the editing network. For editing region reasoning, we train an instruction-based editing region generation network with a multi-modal large language model. Finally, a hint-guided instruction-based editing network is proposed for editing image generations based on the sizeable text-to-image diffusion model to accept the hints for generation. Extensive experiments demonstrate that our method has competitive editing abilities on complex real-world images.

图像编辑多模态指令生成

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。