用衣物为中心的生成方法,精准控制模特穿搭效果。
Fine-Grained Controllable Apparel Showcase Image Generation via Garment-Centric Outpainting
- 以衣物图像为输入,避免学习布料变形,保留细节。
- 支持文本和人脸图像控制,生成多姿态穿搭图。
- 轻量融合模块提升效率,适合电商与设计应用。
本文提出一种基于潜在扩散模型(LDM)的衣物中心外扩(GCO)框架,用于精细化可控的服装展示图像生成。该框架通过给定的衣物图像(来自人台或真人)与文本提示、面部图像,定制化生成穿着特定衣物的模特图像。与现有方法不同,本框架直接以分割出的衣物图像为输入,无需学习布料形变,确保衣物细节忠实还原。框架包含两个阶段:第一阶段采用自适应姿态预测模型,根据衣物生成多样姿态;第二阶段在衣物、预测姿态、文本提示和面部图像条件下生成展示图。特别地,设计了多尺度外观定制模块(MS-ACM),实现整体与局部的文本控制。同时,采用轻量级特征融合操作,无需额外编码器,提升效率。大量实验验证了该框架优于当前最优方法。
原文摘要 · Abstract (English)
In this paper, we propose a novel garment-centric outpainting (GCO) framework based on the latent diffusion model (LDM) for fine-grained controllable apparel showcase image generation. The proposed framework aims at customizing a fashion model wearing a given garment via text prompts and facial images. Different from existing methods, our framework takes a garment image segmented from a dressed mannequin or a person as the input, eliminating the need for learning cloth deformation and ensuring faithful preservation of garment details. The proposed framework consists of two stages. In the first stage, we introduce a garment-adaptive pose prediction model that generates diverse poses given the garment. Then, in the next stage, we generate apparel showcase images, conditioned on the garment and the predicted poses, along with specified text prompts and facial images. Notably, a multi-scale appearance customization module (MS-ACM) is designed to allow both overall and fine-grained text-based control over the generated model's appearance. Moreover, we leverage a lightweight feature fusion operation without introducing any extra encoders or modules to integrate multiple conditions, which is more efficient. Extensive experiments validate the superior performance of our framework compared to state-of-the-art methods.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。