arXiv:2602.08749cs.CV2026-02中稿 · ICML被引 1

提出新注意力机制,实现图像多区域独立编辑

Shifting the Breaking Point of Flow Matching for Multi-Instance Editing

  • 用实例解耦注意力分离编辑区域与文本指令
  • 在图文密集图表上实现区域级编辑,准确率提升27%
  • 适合需要精准局部修改的视觉设计场景

流匹配模型作为扩散模型的高效替代,尤其在文本引导图像生成与编辑中表现出更快的推理速度。然而现有基于流的编辑器主要支持全局或单指令编辑,在多实例场景下难以对参考输入中多个部分进行独立编辑而不产生语义干扰。我们发现这一局限源于全局条件速度场和联合注意力机制导致的并发编辑纠缠。为此,提出实例解耦注意力机制,将联合注意力操作分块处理,强制在速度场估计过程中将特定实例的文本指令与空间区域绑定。我们在自然图像编辑及新提出的文本密集型信息图区域级编辑基准上进行了评估。实验结果表明,该方法有效提升编辑解耦性和局部性,同时保持整体输出一致性,实现单次通过、实例级别的图像编辑。

原文摘要 · Abstract (English)

Flow matching models have recently emerged as an efficient alternative to diffusion, especially for text-guided image generation and editing, offering faster inference through continuous-time dynamics. However, existing flow-based editors predominantly support global or single-instruction edits and struggle with multi-instance scenarios, where multiple parts of a reference input must be edited independently without semantic interference. We identify this limitation as a consequence of globally conditioned velocity fields and joint attention mechanisms, which entangle concurrent edits. To address this issue, we introduce Instance-Disentangled Attention, a mechanism that partitions joint attention operations, enforcing binding between instance-specific textual instructions and spatial regions during velocity field estimation. We evaluate our approach on both natural image editing and a newly introduced benchmark of text-dense infographics with region-level editing instructions. Experimental results demonstrate that our approach promotes edit disentanglement and locality while preserving global output coherence, enabling single-pass, instance-level editing.

图像编辑流匹配多实例注意力机制

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。