通过分步推理实现复杂场景中精准图像编辑
InterCoG: Towards Spatially Precise Image Editing with Interleaved Chain-of-Grounding Reasoning
- 先在文本中分析空间关系定位目标,再用视觉标注确认位置
- 在45K样本数据集上实现高精度编辑,显著优于现有方法
- 适合需要精细空间操作的图像编辑任务
新兴的统一编辑模型在通用物体编辑任务中表现优异,但在复杂多实体场景中进行细粒度编辑仍具挑战,尤其当目标不显眼且需空间推理时。为此,我们提出InterCoG,一种基于文本-视觉交错式链式接地推理的精细图像编辑框架。其核心思路是:首先仅通过包含空间关系的文本推理目标的位置与身份;随后在像素空间中生成边界框和掩码完成视觉定位;最后重写编辑描述以明确预期结果。为支持该范式,我们设计了两个辅助训练模块:多模态接地重建监督(强化空间定位精度)和多模态接地推理对齐(提升推理可解释性)。同时构建了GroundEdit-45K数据集(含45,000个带详细推理标注的编辑样本)及GroundEdit-Bench评估基准。大量实验证明,本方法在空间复杂、多实体场景下具有显著更高的编辑精度。
原文摘要 · Abstract (English)
Emerging unified editing models have demonstrated strong capabilities in general object editing tasks. However, it remains a significant challenge to perform fine-grained editing in complex multi-entity scenes, particularly those where targets are not visually salient and require spatial reasoning. To this end, we propose InterCoG, a novel text-vision Interleaved Chain-of-Grounding reasoning framework for fine-grained image editing in complex real-world scenes. The key insight of InterCoG is to first perform object position reasoning solely within text that includes spatial relation details to explicitly deduce the location and identity of the edited target. It then conducts visual grounding via highlighting the editing targets with generated bounding boxes and masks in pixel space, and finally rewrites the editing description to specify the intended outcomes. To further facilitate this paradigm, we propose two auxiliary training modules: multimodal grounding reconstruction supervision and multimodal grounding reasoning alignment to enforce spatial localization accuracy and reasoning interpretability, respectively. We also construct GroundEdit-45K, a dataset comprising 45K grounding-oriented editing samples with detailed reasoning annotations, and GroundEdit-Bench for grounding-aware editing evaluation. Extensive experiments substantiate the superiority of our approach in highly precise edits under spatially intricate and multi-entity scenes.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。