arXiv:2601.03741cs.CV2026-01ACL被引 3

将图像转为可操作对象层,实现精准文本编辑

I2E: From Image Pixels to Actionable Interactive Environments for Text-Guided Image Editing

  • 分步执行:先拆解图像为可操作对象,再按指令执行原子动作
  • 在多对象空间推理任务中优于现有方法,保持物理合理性
  • 适合需要精细控制和复杂编辑的用户

现有文本引导图像编辑方法主要依赖端到端像素级修复范式。尽管在简单场景中表现良好,但在需要精确局部控制和复杂多对象空间推理的组合编辑任务中仍面临挑战。该范式受限于1)规划与执行的隐式耦合,2)缺乏对象级控制粒度,3)依赖非结构化的像素中心建模。为此,我们提出I2E,一种全新的“分解-执行”范式,将图像编辑重构为结构化环境中的可行动作过程。I2E通过分解器将非结构化图像转换为离散、可操作的对象图层,并引入具备物理感知能力的视觉-语言-动作智能体,通过思维链推理将复杂指令解析为一系列原子动作。此外,我们还构建了I2E-Bench,一个专用于多实例空间推理与高精度编辑的基准测试集。在I2E-Bench及多个公开基准上的实验结果表明,I2E在处理复杂组合指令、保持物理合理性以及确保多轮编辑稳定性方面显著优于现有最先进方法。

原文摘要 · Abstract (English)

Existing text-guided image editing methods primarily rely on end-to-end pixel-level inpainting paradigm. Despite its success in simple scenarios, this paradigm still significantly struggles with compositional editing tasks that require precise local control and complex multi-object spatial reasoning. This paradigm is severely limited by 1) the implicit coupling of planning and execution, 2) the lack of object-level control granularity, and 3) the reliance on unstructured, pixel-centric modeling. To address these limitations, we propose I2E, a novel "Decompose-then-Action" paradigm that revisits image editing as an actionable interaction process within a structured environment. I2E utilizes a Decomposer to transform unstructured images into discrete, manipulable object layers and then introduces a physics-aware Vision-Language-Action Agent to parse complex instructions into a series of atomic actions via Chain-of-Thought reasoning. Further, we also construct I2E-Bench, a benchmark designed for multi-instance spatial reasoning and high-precision editing. Experimental results on I2E-Bench and multiple public benchmarks demonstrate that I2E significantly outperforms state-of-the-art methods in handling complex compositional instructions, maintaining physical plausibility, and ensuring multi-turn editing stability.

图像编辑多对象推理文本生成

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。