arXiv:2603.01893cs.CV2026-03被引 3

让AI像人一样思考再编辑图像,提升复杂场景下的精准定位能力。

Generative Visual Chain-of-Thought for Image Editing

  • 先生成空间线索定位目标区域,再执行编辑,端到端优化视觉推理与操作
  • 在19个任务上训练180万样本,性能超越现有最佳模型
  • 适合需要高精度、可解释性图像编辑的研究与应用

现有图像编辑方法在复杂场景和细微空间指令下难以准确定位编辑区域。为此,我们提出生成式视觉思维链(GVCoT),通过先生成空间线索定位目标区域,再执行编辑的统一框架,实现原生视觉推理。与仅依赖文本或工具的视觉思维链不同,GVCoT在推理与编辑阶段联合优化生成的视觉标记,端到端训练中自发形成空间推理能力,更有效利用视觉线索。训练难点在于缺乏带精确编辑区域标注的大规模数据,因此我们构建了包含180万高质量样本的GVCoT-Edit-Instruct数据集,覆盖19个任务。采用渐进式训练策略:先通过监督微调建立推理轨迹中的定位基础,再通过强化学习提升推理与编辑质量。最后,我们设计了SREdit-Bench新基准,用于在复杂场景和细粒度指代表达下全面评估模型表现。实验表明,GVCoT在SREdit-Bench和ImgEdit上持续优于当前最优模型。我们希望该工作能推动可解释且精准的图像编辑研究发展。

原文摘要 · Abstract (English)

Existing image editing methods struggle to perceive where to edit, especially under complex scenes and nuanced spatial instructions. To address this issue, we propose Generative Visual Chain-of-Thought (GVCoT), a unified framework that performs native visual reasoning by first generating spatial cues to localize the target region and then executing the edit. Unlike prior text-only CoT or tool-dependent visual CoT paradigms, GVCoT jointly optimizes visual tokens generated during the reasoning and editing phases in an end-to-end manner. This way fosters the emergence of innate spatial reasoning ability and enables more effective utilization of visual-domain cues. The main challenge of training GCVoT lies in the scarcity of large-scale editing data with precise edit region annotations; to this end, we construct GVCoT-Edit-Instruct, a dataset of 1.8M high-quality samples spanning 19 tasks. We adopt a progressive training strategy: supervised fine-tuning to build foundational localization ability in reasoning trace before final editing, followed by reinforcement learning to further improve reasoning and editing quality. Finally, we introduce SREdit-Bench, a new benchmark designed to comprehensively stress-test models under sophisticated scenes and fine-grained referring expressions. Experiments demonstrate that GVCoT consistently outperforms state-of-the-art models on SREdit-Bench and ImgEdit. We hope our GVCoT will inspire future research toward interpretable and precise image editing.

图像编辑视觉推理生成模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。