一键生成穿搭图:支持多服饰、人脸、姿势自定义。
FashionComposer: Compositional Fashion Image Generation
- 用多模态输入统一框架,灵活组合文本、人体模型和参考图。
- 通过资产库与绑定注意力机制,实现多参考图精准融合。
- 适合虚拟试衣、人物合集生成等创意设计场景。
我们提出 FashionComposer,用于组合式服装图像生成。与以往方法不同,FashionComposer 具有高度灵活性,可接受多模态输入(文本提示、参数化人体模型、服装图像和人脸图像),并支持在一次推理中个性化人体外观、姿态与体型,同时分配多个服装。为此,我们首先构建一个能处理多种输入模态的通用框架,并通过缩放训练数据增强模型的组合能力。为无缝整合多个参考图像(服装和人脸),我们将这些参考图像组织成单张“资产库”图像,并使用参考UNet提取外观特征。为将外观特征准确注入生成结果的对应像素,我们提出主体绑定注意力(subject-binding attention),将不同“资产”的特征与相应文本特征绑定。由此,模型可根据语义理解每个资产,支持任意数量和类型的参考图像。作为综合性解决方案,FashionComposer 还支持人像合集生成、多样虚拟试衣等应用。
原文摘要 · Abstract (English)
We present FashionComposer for compositional fashion image generation. Unlike previous methods, FashionComposer is highly flexible. It takes multi-modal input (i.e., text prompt, parametric human model, garment image, and face image) and supports personalizing the appearance, pose, and figure of the human and assigning multiple garments in one pass. To achieve this, we first develop a universal framework capable of handling diverse input modalities. We construct scaled training data to enhance the model's robust compositional capabilities. To accommodate multiple reference images (garments and faces) seamlessly, we organize these references in a single image as an "asset library" and employ a reference UNet to extract appearance features. To inject the appearance features into the correct pixels in the generated result, we propose subject-binding attention. It binds the appearance features from different "assets" with the corresponding text features. In this way, the model could understand each asset according to their semantics, supporting arbitrary numbers and types of reference images. As a comprehensive solution, FashionComposer also supports many other applications like human album generation, diverse virtual try-on tasks, etc.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。