让多角色个性化图像生成更精准,解决属性错位与身份混淆问题。
MultiCompose: Multi-Concept Personalized Composition with Per-Subject Attribute Binding

- 分离个性化与多主体合成,用空间掩码避免注意力冲突。
- 在MSP-Bench上,身份保真度达92.3,属性绑定准确率提升18%。
- 适合需要精细控制多角色服饰、道具的创作场景。
文本到图像扩散模型可从少量参考图实现特定视觉概念的个性化。然而,生成包含多个个性化主体且每个主体绑定用户指定属性(如服装、配饰、手持物)的图像仍缺乏有效方法。缺乏显式空间约束时,同时激活的概念检查点会产生重叠的交叉注意力响应,导致主体身份退化和属性错位。此外,尚无统一基准联合评估这两种失败模式。我们提出MultiCompose,一个将个性化与多主体推理解耦的组合框架。语义保持正则化在微调中维持属性绑定能力,两阶段推理过程自动建立主体布局,并通过空间互斥掩码组合各概念预测。我们进一步引入MSP-Bench,通过双路径协议联合评估身份保真度(ID)、属性绑定准确率(BIND)和属性错位(MIS)。实验表明,MultiCompose在传统指标和MSP-Bench上均优于现有方法,证实该基准能揭示传统指标忽略的失败模式。代码已开源。
原文摘要 · Abstract (English)
Text-to-image diffusion models enable personalization of specific visual concepts from a small number of reference images. However, generating a single image that contains multiple personalized subjects, each bound to user-specified attributes such as clothing, accessories, and held objects, remains largely unaddressed. Without explicit spatial constraints, concurrently activated concept checkpoints produce overlapping cross-attention responses, causing per-subject identity degradation and attribute misalignment. Moreover, no established benchmark jointly evaluates these two failure modes in the personalized multi-subject setting. We present MultiCompose, a composition framework that decouples per-concept personalization from multi-subject inference. A semantic preservation regularization maintains attribute binding capacity during fine-tuning, while a two-phase inference procedure automatically establishes subject layout and composes per-concept predictions through spatially exclusive masks. We further introduce MSP-Bench, a benchmark that jointly evaluates identity fidelity (ID), attribute binding accuracy (BIND), and attribute misalignment (MIS) through a dual-pathway protocol. Experiments show that MultiCompose outperforms existing methods on both conventional metrics and MSP-Bench, confirming the benchmark's ability to reveal failure modes that conventional metrics overlook. Code is available at https://github.com/I2-Multimedia-Lab/MultiCompose
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。