轻量级服装生成模型,能精准适配各种服饰与提示。
DreamFit: Garment-Centric Human Generation via a Lightweight Anything-Dressing Encoder
- 用自适应注意力和LoRA模块降低参数至8340万,训练轻量。
- 支持任意服装类型与创意风格,生成效果稳定高质量。
- 可无缝接入现有扩散模型插件,适合快速部署使用。
面向文本或图像提示的服装中心人体生成扩散模型受到广泛关注,具备广泛应用潜力。然而,现有方法面临两难:轻量级方案(如适配器)易生成不一致纹理;微调类方法训练成本高,且难以保持预训练扩散模型的泛化能力,限制其在多样化场景中的表现。为此,我们提出DreamFit,采用专为服装中心人体生成设计的轻量级Anything-Dressing编码器。该模型具三大优势:(1) 轻量训练:通过自适应注意力与LoRA模块,将模型参数压缩至8340万;(2) 万物适配:对各类(非)服装、创意风格及提示指令均有良好泛化性,跨场景生成质量稳定;(3) 即插即用:可与任意社区控制插件无缝集成,降低使用门槛。为进一步提升生成质量,DreamFit利用预训练多模态大模型(LMMs)增强提示,提供细粒度服装描述,缩小训练与推理间的提示差距。我们在768×512高分辨率基准和真实场景图像上进行了全面实验,结果表明DreamFit超越所有现有方法,展现出领先的服装中心人体生成能力。
原文摘要 · Abstract (English)
Diffusion models for garment-centric human generation from text or image prompts have garnered emerging attention for their great application potential. However, existing methods often face a dilemma: lightweight approaches, such as adapters, are prone to generate inconsistent textures; while finetune-based methods involve high training costs and struggle to maintain the generalization capabilities of pretrained diffusion models, limiting their performance across diverse scenarios. To address these challenges, we propose DreamFit, which incorporates a lightweight Anything-Dressing Encoder specifically tailored for the garment-centric human generation. DreamFit has three key advantages: (1) \textbf{Lightweight training}: with the proposed adaptive attention and LoRA modules, DreamFit significantly minimizes the model complexity to 83.4M trainable parameters. (2)\textbf{Anything-Dressing}: Our model generalizes surprisingly well to a wide range of (non-)garments, creative styles, and prompt instructions, consistently delivering high-quality results across diverse scenarios. (3) \textbf{Plug-and-play}: DreamFit is engineered for smooth integration with any community control plugins for diffusion models, ensuring easy compatibility and minimizing adoption barriers. To further enhance generation quality, DreamFit leverages pretrained large multi-modal models (LMMs) to enrich the prompt with fine-grained garment descriptions, thereby reducing the prompt gap between training and inference. We conduct comprehensive experiments on both $768 \times 512$ high-resolution benchmarks and in-the-wild images. DreamFit surpasses all existing methods, highlighting its state-of-the-art capabilities of garment-centric human generation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。