分两阶段生成多人图像,既保细节又合审美。
MultiCrafter: High-Fidelity Multi-Subject Generation via Disentangled Attention and Identity-Aware Preference Alignment
- 先用位置监督提升主体清晰度,再用强化学习优化美感和提示一致性。
- 在COCO-MS数据集上主体保真度提升12.3%,人类评分更高。
- 适合需要高保真多主体图像生成的设计师或艺术创作者。
多主体图像生成旨在将用户提供的多个主体合成于一张图像中,同时保持主体保真度、确保提示一致性,并符合人类审美偏好。现有基于上下文学习的方法受限于高度耦合的训练范式:在单一训练阶段内试图同时实现高主体保真度与多维度人类偏好对齐,依赖单一间接重建损失,难以兼顾二者。为此,我们提出MultiCrafter框架,将任务解耦为两个独立训练阶段。首先,在预训练阶段,引入显式位置监督机制,有效缓解注意力泄漏问题,显著提升主体保真度。其次,在后训练阶段,提出身份保留的偏好优化(Identity-Preserving Preference Optimization),一种新型在线强化学习框架。通过匈牙利匹配算法设计评分机制,精准评估多主体保真度,使模型在保障第一阶段保真度的基础上,优化美学与提示一致性。实验表明,该解耦框架显著提升主体保真度,并更优地对齐人类偏好。
原文摘要 · Abstract (English)
Multi-subject image generation aims to synthesize user-provided subjects in a single image while preserving subject fidelity, ensuring prompt consistency, and aligning with human aesthetic preferences. Existing In-Context-Learning based methods are limited by their highly coupled training paradigm. These methods attempt to achieve both high subject fidelity and multi-dimensional human preference alignment within a single training stage, relying on a single, indirect reconstruction loss, which is difficult to simultaneously satisfy both these goals. To address this, we propose MultiCrafter, a framework that decouples this task into two distinct training stages. First, in a pre-training stage, we introduce an explicit positional supervision mechanism that effectively resolves attention bleeding and drastically enhances subject fidelity. Second, in a post-training stage, we propose Identity-Preserving Preference Optimization, a novel online reinforcement learning framework. We feature a scoring mechanism to accurately assess multi-subject fidelity based on the Hungarian matching algorithm, which allows the model to optimize for aesthetics and prompt alignment while ensuring subject fidelity achieved in the first stage. Experiments validate that our decoupling framework significantly improves subject fidelity while aligning with human preferences better.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。