arXiv:2511.21087cs.CV2025-11被引 8

MIRA通过迭代推理让AI更准确理解复杂图像编辑指令。

MIRA: Multimodal Iterative Reasoning Agent for Image Editing

  • 采用视觉反馈驱动的多轮推理机制,逐步生成编辑指令。
  • 在150K数据集上训练,显著提升语义一致性和视觉质量。
  • 可插拔适配开源模型,效果媲美甚至超越闭源系统。

基于指令的图像编辑为用户提供了自然语言操作图像的直观方式。然而,扩散模型在处理涉及组合关系、上下文线索或指代表达的复杂指令时,常因理解偏差导致语义漂移或编辑失败。为此,本文提出轻量级、可插拔的多模态迭代推理代理MIRA(Multimodal Iterative Reasoning Agent),通过感知-推理-行动的循环模拟多轮人机交互过程。MIRA不依赖单次提示或静态计划,而是基于视觉反馈逐步预测原子化编辑指令。结合自建的150K多模态工具使用数据集MIRA-Editing及两阶段SFT+GRPO训练流程,使MIRA能有效处理复杂指令。当与Flux.1-Kontext、Step1X-Edit、Qwen-Image-Edit等开源编辑模型结合时,显著提升语义一致性与感知质量,性能达到甚至超过GPT-Image、Nano-Banana等闭源系统。

原文摘要 · Abstract (English)

Instruction-guided image editing offers an intuitive way for users to edit images with natural language. However, diffusion-based editing models often struggle to accurately interpret complex user instructions, especially those involving compositional relationships, contextual cues, or referring expressions, leading to edits that drift semantically or fail to reflect the intended changes. We tackle this problem by proposing MIRA (Multimodal Iterative Reasoning Agent), a lightweight, plug-and-play multimodal reasoning agent that performs editing through an iterative perception-reasoning-action loop, effectively simulating multi-turn human-model interaction processes. Instead of issuing a single prompt or static plan, MIRA predicts atomic edit instructions step by step, using visual feedback to make its decisions. Our 150K multimodal tool-use dataset, MIRA-Editing, combined with a two-stage SFT + GRPO training pipeline, enables MIRA to perform reasoning and editing over complex editing instructions. When paired with open-source image editing models such as Flux.1-Kontext, Step1X-Edit, and Qwen-Image-Edit, MIRA significantly improves both semantic consistency and perceptual quality, achieving performance comparable to or exceeding proprietary systems such as GPT-Image and Nano-Banana.

图像编辑多模态推理扩散模型自然语言控制

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。