arXiv:2412.04146cs.CV2024-12CVPR被引 10

用扩散模型实现任意组合服装的精准虚拟试穿。

AnyDressing: Customizable Multi-Garment Virtual Dressing via Latent Diffusion Models

  • 分步提取服装纹理,避免混淆,提升细节保留。
  • 多服装区域精准注入,文本与图像一致性更强。
  • 可插件式接入主流生成模型,适合内容创作者使用。

基于扩散模型的服饰生成技术虽有进展,但现有方法难以支持多种服装组合,且在保持文本提示忠实度的同时难以保留服装细节。本文提出AnyDressing,一种面向任意服装组合与个性化文本提示的多服装虚拟试穿方法。该方法包含两个核心网络:用于提取服装特征的GarmentsNet和用于生成图像的DressingNet。其中,GarmentsNet引入服装特异性特征提取模块,可并行编码各类服装纹理,避免混淆同时保持高效。DressingNet设计了自适应的Dressing-Attention机制与实例级服装定位学习策略,精确将多服装特征注入对应区域,增强纹理融合与文本一致性。此外,通过服装增强纹理学习策略进一步优化细粒度纹理表现。实验表明,AnyDressing在多个数据集上达到当前最优性能,并可作为插件轻松集成至任意扩散模型控制扩展中,显著提升生成图像的多样性与可控性。

原文摘要 · Abstract (English)

Recent advances in garment-centric image generation from text and image prompts based on diffusion models are impressive. However, existing methods lack support for various combinations of attire, and struggle to preserve the garment details while maintaining faithfulness to the text prompts, limiting their performance across diverse scenarios. In this paper, we focus on a new task, i.e., Multi-Garment Virtual Dressing, and we propose a novel AnyDressing method for customizing characters conditioned on any combination of garments and any personalized text prompts. AnyDressing comprises two primary networks named GarmentsNet and DressingNet, which are respectively dedicated to extracting detailed clothing features and generating customized images. Specifically, we propose an efficient and scalable module called Garment-Specific Feature Extractor in GarmentsNet to individually encode garment textures in parallel. This design prevents garment confusion while ensuring network efficiency. Meanwhile, we design an adaptive Dressing-Attention mechanism and a novel Instance-Level Garment Localization Learning strategy in DressingNet to accurately inject multi-garment features into their corresponding regions. This approach efficiently integrates multi-garment texture cues into generated images and further enhances text-image consistency. Additionally, we introduce a Garment-Enhanced Texture Learning strategy to improve the fine-grained texture details of garments. Thanks to our well-craft design, AnyDressing can serve as a plug-in module to easily integrate with any community control extensions for diffusion models, improving the diversity and controllability of synthesized images. Extensive experiments show that AnyDressing achieves state-of-the-art results.

虚拟试穿扩散模型服装生成

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。