arXiv:2509.18092cs.CV2025-09SIGGRAPH被引 10

用不同参考图控制人物发型、衣着等细节,实现精准可控的人像生成。

ComposeMe: Attribute-Specific Image Prompts for Controllable Human Image Generation

  • 分属性输入参考图,生成对应特征的图像令牌注入扩散模型。
  • 在多属性组合下仍保持身份一致与文本提示准确,效果领先现有方法。
  • 适合需要精细调节人物外观的研究者或创意设计用户。

在个性化文生图任务中,对发型、衣着等细粒度属性实现高保真且可控的人像生成仍是核心挑战。现有方法虽注重从参考图保留身份信息,但缺乏模块化设计,难以解耦控制特定视觉属性。本文提出一种新型属性特定图像提示范式:使用独立的参考图集分别引导头发、服装和身份等外观要素的生成。将这些输入编码为属性专属令牌,并注入预训练文生图扩散模型中,实现多个视觉因素的组合与解耦控制,甚至可在单张图像中处理多人。为提升自然融合与鲁棒解耦能力,我们构建了一个包含多样化姿态与表情的跨参考训练数据集,并提出多属性跨参考训练策略,使模型在属性输入错位时仍能忠实生成符合身份与文本条件的结果。大量实验表明,本方法在遵循视觉与文本提示方面达到当前最优性能。该框架通过结合视觉提示与文本驱动生成,为更灵活可配置的人像合成开辟了新路径。官网:https://snap-research.github.io/composeme/

原文摘要 · Abstract (English)

Generating high-fidelity images of humans with fine-grained control over attributes such as hairstyle and clothing remains a core challenge in personalized text-to-image synthesis. While prior methods emphasize identity preservation from a reference image, they lack modularity and fail to provide disentangled control over specific visual attributes. We introduce a new paradigm for attribute-specific image prompting, in which distinct sets of reference images are used to guide the generation of individual aspects of human appearance, such as hair, clothing, and identity. Our method encodes these inputs into attribute-specific tokens, which are injected into a pre-trained text-to-image diffusion model. This enables compositional and disentangled control over multiple visual factors, even across multiple people within a single image. To promote natural composition and robust disentanglement, we curate a cross-reference training dataset featuring subjects in diverse poses and expressions, and propose a multi-attribute cross-reference training strategy that encourages the model to generate faithful outputs from misaligned attribute inputs while adhering to both identity and textual conditioning. Extensive experiments show that our method achieves state-of-the-art performance in accurately following both visual and textual prompts. Our framework paves the way for more configurable human image synthesis by combining visual prompting with text-driven generation. Webpage is available at: https://snap-research.github.io/composeme/.

人像生成属性控制扩散模型视觉提示

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。