arXiv:2410.12501cs.CVcs.AI2024-10被引 4

用文本驱动的注意力机制,让虚拟试穿更精准还原服装细节。

DH-VTON: Deep Text-Driven Virtual Try-On via Hybrid Attention Learning

  • 引入InternViT-6B提取服装深层语义特征,增强文本对齐能力。
  • 提出混合注意力策略,多尺度保留服装纹理与结构细节。
  • 适合关注服装细节还原的生成模型研究者与电商应用开发者。

虚拟试穿(VTON)旨在合成特定人物穿着指定衣物的图像,近年来在在线购物场景中受到广泛关注。当前核心挑战在于深度估计时对参考衣物的细粒度语义提取(即深层语义)以及将衣物合成并形变到人体时的有效纹理保持。为此,我们提出DH-VTON,一种基于深度文本驱动的虚拟试穿模型,包含独特的混合注意力学习策略和深层服装语义保持模块。依托成熟的预训练画图示例(PBE)框架,本工作构建了完整流程:首先引入InternViT-6B作为细粒度特征学习器,可与大规模内在知识及深层文本语义(如“领口”或“腰带”)对齐,弥补常用CLIP编码器的不足;在此基础上,为增强定制化穿搭能力,进一步引入服装特征控制网升级版(GFC+)模块,并提出一种新型混合注意力训练策略,能自适应地将服装的细粒度特征融入VTON模型不同层级,实现多尺度特征保留。在多个代表性数据集上的大量实验表明,该方法优于以往基于扩散模型和基于GAN的方法,在保持服装细节和生成真实人体图像方面表现优异。

原文摘要 · Abstract (English)

Virtual Try-ON (VTON) aims to synthesis specific person images dressed in given garments, which recently receives numerous attention in online shopping scenarios. Currently, the core challenges of the VTON task mainly lie in the fine-grained semantic extraction (i.e.,deep semantics) of the given reference garments during depth estimation and effective texture preservation when the garments are synthesized and warped onto human body. To cope with these issues, we propose DH-VTON, a deep text-driven virtual try-on model featuring a special hybrid attention learning strategy and deep garment semantic preservation module. By standing on the shoulder of a well-built pre-trained paint-by-example (abbr. PBE) approach, we present our DH-VTON pipeline in this work. Specifically, to extract the deep semantics of the garments, we first introduce InternViT-6B as fine-grained feature learner, which can be trained to align with the large-scale intrinsic knowledge with deep text semantics (e.g.,"neckline" or "girdle") to make up for the deficiency of the commonly adopted CLIP encoder. Based on this, to enhance the customized dressing abilities, we further introduce Garment-Feature ControlNet Plus (abbr. GFC+) module and propose to leverage a fresh hybrid attention strategy for training, which can adaptively integrate fine-grained characteristics of the garments into the different layers of the VTON model, so as to achieve multi-scale features preservation effects. Extensive experiments on several representative datasets demonstrate that our method outperforms previous diffusion-based and GAN-based approaches, showing competitive performance in preserving garment details and generating authentic human images.

虚拟试穿文本驱动注意力机制服装生成

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。