arXiv:2507.10217cs.CV2025-07

用少量衣服图就能精准生成人物全身照,还能按提示换装。

From Wardrobe to Canvas: Wardrobe Polyptych LoRA for Part-level Controllable Human Image Generation

  • 仅训练LoRA层,推理时无额外参数负担。
  • 在新主体上生成图像时保持身份与服装细节高度一致。
  • 适合需要快速个性化生成的实时应用,如虚拟试衣。

近期扩散模型通过学习特定主体实现个性化,使所学属性融入生成图像中。然而,个性化人体图像生成仍面临挑战,需精确且一致地保留属性(如身份、衣物细节)。现有方法要么需在推理时用少量图像微调,要么依赖大规模数据集训练,二者均计算开销大,难以用于实时场景。为此,我们提出Wardrobe Polyptych LoRA,一种面向局部可控的个性化人体图像生成新方法。通过仅训练LoRA层,该方法在推理阶段无计算负担,同时实现未见主体的高保真合成。核心思想是将生成条件置于主体的衣橱信息,并利用空间参考减少信息丢失,从而提升保真度与一致性。此外,引入选择性主体区域损失,使模型在训练中忽略部分参考图像,增强生成结果与文本提示的对齐性,同时维持主体完整性。显著特点是推理无需额外参数,仅用一个在少量样本上训练的模型即可完成生成。我们构建了专用于个性化人体图像生成的新数据集与基准。大量实验表明,本方法在保真度与一致性上显著优于现有技术,支持真实且身份一致的全身合成。

原文摘要 · Abstract (English)

Recent diffusion models achieve personalization by learning specific subjects, allowing learned attributes to be integrated into generated images. However, personalized human image generation remains challenging due to the need for precise and consistent attribute preservation (e.g., identity, clothing details). Existing subject-driven image generation methods often require either (1) inference-time fine-tuning with few images for each new subject or (2) large-scale dataset training for generalization. Both approaches are computationally expensive and impractical for real-time applications. To address these limitations, we present Wardrobe Polyptych LoRA, a novel part-level controllable model for personalized human image generation. By training only LoRA layers, our method removes the computational burden at inference while ensuring high-fidelity synthesis of unseen subjects. Our key idea is to condition the generation on the subject's wardrobe and leverage spatial references to reduce information loss, thereby improving fidelity and consistency. Additionally, we introduce a selective subject region loss, which encourages the model to disregard some of reference images during training. Our loss ensures that generated images better align with text prompts while maintaining subject integrity. Notably, our Wardrobe Polyptych LoRA requires no additional parameters at the inference stage and performs generation using a single model trained on a few training samples. We construct a new dataset and benchmark tailored for personalized human image generation. Extensive experiments show that our approach significantly outperforms existing techniques in fidelity and consistency, enabling realistic and identity-preserving full-body synthesis.

个性化生成可控生成LoRA

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。