解决多主体图像定制中的位置与身份一致性问题,实现精准可控的高保真生成。
PositionIC: Unified Position and Identity Consistency for Image Customization
- 构建首个自动合成带位置标注的多主体数据集,提供关键空间监督
- 提出可见性感知注意力机制,解耦空间布局与身份特征,支持遮挡感知放置
- 在公开基准上刷新空间精度与身份一致性记录,适合多主体图像编辑场景
近期基于主体的图像定制在保真度方面表现优异,但细粒度实例级空间控制仍是未解难题,限制了实际应用。这一瓶颈源于两个因素:缺乏可扩展的位置标注数据集,以及全局注意力机制导致身份与布局信息纠缠。为此,我们提出PositionIC,一种统一的高保真、空间可控的多主体定制框架。首先,提出BMPDS,首个自动生成位置标注的多主体数据集合成管道,有效提供关键空间监督。其次,设计轻量级、布局感知的扩散框架,引入新型可见性感知注意力机制。该机制通过受NeRF启发的体素权重调节,显式建模空间关系,有效解耦实例级空间嵌入与语义身份特征,实现精确、遮挡感知的多主体放置。大量实验表明,PositionIC在公共基准上达到当前最佳性能,创下空间精度与身份一致性的新纪录。本工作为多实体场景下的真正可控高保真图像定制迈出了重要一步。代码与数据:https://github.com/MeiGen-AI/PositionIC。
原文摘要 · Abstract (English)
Recent subject-driven image customization excels in fidelity, yet fine-grained instance-level spatial control remains an elusive challenge, hindering real-world applications. This limitation stems from two factors: a scarcity of scalable, position-annotated datasets, and the entanglement of identity and layout by global attention mechanisms. To this end, we introduce PositionIC, a unified framework for high-fidelity, spatially controllable multi-subject customization. First, we present BMPDS, the first automatic data-synthesis pipeline for position-annotated multi-subject datasets, effectively providing crucial spatial supervision. Second, we design a lightweight, layout-aware diffusion framework that integrates a novel visibility-aware attention mechanism. This mechanism explicitly models spatial relationships via an NeRF-inspired volumetric weight regulation to effectively decouple instance-level spatial embeddings from semantic identity features, enabling precise, occlusion-aware placement of multiple subjects. Extensive experiments demonstrate PositionIC achieves state-of-the-art performance on public benchmarks, setting new records for spatial precision and identity consistency. Our work represents a significant step towards truly controllable, high-fidelity image customization in multi-entity scenarios. Code and data: https://github.com/MeiGen-AI/PositionIC.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。