让图像编辑保持物体数量和位置不变,同时改换物体语义。
IMAGHarmony: Controllable Image Editing with Consistent Object Quantity and Layout
- 引入感知提示模块,联合推理物体语义、数量与位置。
- 在多物体场景下保持结构一致,错误率降低40%以上。
- 仅需200张训练图,适合快速部署的可控编辑任务。
尽管基于扩散模型的图像编辑取得进展,多物体场景的操控仍具挑战。现有方法常以牺牲结构一致性为代价实现语义修改,无法在不引入意外移动或背景变化的前提下精确保持物体数量与空间布局。为此,我们提出量与布局一致的图像编辑(QL-Edit),在改变物体语义的同时保持原始实例数量与空间布局。提出IMAGHarmony框架,包含感知一致(HA)模块,将参考图像的感知线索融入扩散过程,使模型能联合推理物体语义、数量与空间位置,提升结构一致性。此外,设计偏好引导噪声选择(PNS)策略,识别有利初始化条件,在复杂多物体场景中显著提升生成稳定性。为支持系统评估,构建HarmonyBench基准,用于衡量在数量与布局约束下的语义编辑准确率与结构一致性。大量实验表明,IMAGHarmony在结构保留与语义准确性上均优于现有方法。值得注意的是,该框架高效,仅需200张训练图像和1060万可训练参数。代码、模型与数据已公开于\url{https://github.com/muzishen/IMAGHarmony}。
原文摘要 · Abstract (English)
Despite advances in diffusion-based image editing, manipulating multi-object scenes remains challenging. Existing approaches often achieve semantic changes at the expense of structural consistency, failing to preserve exact object counts and spatial layouts without introducing unintended relocations or background modifications. To address this limitation, we introduce quantity-and-layout-consistent image editing (QL-Edit) to modify object semantics while maintaining the original instance cardinality and spatial layout. We propose IMAGHarmony, a parameter-efficient framework featuring a harmony-aware (HA) module that incorporates perception cues from the reference image into the diffusion process. This enables the model to jointly reason about object semantics, counts, and spatial positions for improved structural consistency. Furthermore, we introduce a preference-guided noise selection (PNS) strategy that identifies favorable initialization conditions, substantially improving generation stability in challenging multi-object scenarios. To support systematic evaluation, we construct HarmonyBench, a benchmark designed to measure semantic editing accuracy and structural consistency under quantity and layout constraints. Extensive experiments demonstrate that IMAGHarmony consistently outperforms existing methods in both structural preservation and semantic accuracy. Notably, our framework is highly efficient, requiring only 200 training images and 10.6M trainable parameters. Code, models, and data are available at \url{https://github.com/muzishen/IMAGHarmony}.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。