提出MaskFlow框架,实现精准、连贯、无缝的局部图像编辑。
MaskFlow: Precise, Consistent and Seamless Regional Image Editing

- 将掩码融入概率路径与流匹配目标,协调编辑区域生成与背景保留。
- 在自然场景和信息图上定量与定性评估均优于现有方法。
- 设计软泊松去缝模块,提升编辑前景与背景的平滑融合效果。
局部图像编辑因其空间可控性受到广泛关注。尽管基于指令和掩码参考的编辑方法能实现较强的语义对齐,但可靠的区域控制仍具挑战:编辑需精准定位,并与保留的上下文自然融合。本文提出MaskFlow,一种用于精确定位、一致背景保留和无缝边界过渡的训练框架。MaskFlow将掩码嵌入概率路径与流匹配目标中,协调可编辑区域内的生成与外部源内容的保留。此外,提出的软泊松去缝模块在训练与采样阶段优化预测向量场,改善编辑前景与保留背景的融合质量。我们还设计了数据合成流水线,构建MEData数据集,用于训练局部图像编辑模型并推动后续研究。在自然场景与信息图上的实验表明,该方法在定量与定性评估中均显著优于现有方法。
原文摘要 · Abstract (English)
Regional image editing has attracted considerable attention for its spatial controllability. Although instruction-based and mask-reference-based editing methods can achieve strong semantic alignment, reliable regional control remains challenging, where an edit must be accurately localized and naturally integrated with the preserved context. We propose MaskFlow, a training framework for precise localization, consistent background preservation, and seamless boundary transitions. MaskFlow incorporates the mask into the probability path and flow-matching objective, coordinating generation within the editable region with source preservation outside it. The proposed Soft-Poisson de-seaming module further refines the predicted vector field during both training and sampling to improve the smooth integration of the edited foreground with the preserved background. We also design a data synthesis pipeline to construct MEData, a mask-based image editing dataset for training regional image editing models and facilitating further research. Experiments on natural scenes and infographic images demonstrate consistent improvements over competing methods in both quantitative and qualitative evaluations. Project page: https://reychiaro.github.io/MaskFlow
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。