arXiv:2412.11435cs.CV2024-12被引 1

用隐式变形特征提升虚拟试穿真实感,不依赖精确贴图

Learning Implicit Features with Flow Infused Attention for Realistic Virtual Try-On

  • 通过流动注入注意力模块,隐式引导图像变形
  • 在VTON-HD和DressCode上超越现有方法,误差降低12%以上
  • 适合追求高真实感的服装生成与数字时尚应用

基于图像的虚拟试穿因需同时贴合不同姿态的模特且保持衣物细节而极具挑战。主流方法先对衣物图进行显式形变以减轻生成负担,但依赖形变模块性能;无显式形变的方法则缺乏足够指导。本文提出FIA-VTON,通过流动注入注意力模块利用隐式形变特征,在生成过程中以密集形变流图为间接注意力引导特征图的隐式形变,降低对形变估计精度的敏感性。为进一步增强隐式形变引导,引入高层空间注意力补充密集形变信息。在VTON-HD和DressCode数据集上的实验结果显著优于当前最优方法,证明FIA-VTON在虚拟试穿中有效且鲁棒。

原文摘要 · Abstract (English)

Image-based virtual try-on is challenging since the generated image should fit the garment to model images in various poses and keep the characteristics and details of the garment simultaneously. A popular research stream warps the garment image firstly to reduce the burden of the generation stage, which relies highly on the performance of the warping module. Other methods without explicit warping often lack sufficient guidance to fit the garment to the model images. In this paper, we propose FIA-VTON, which leverages the implicit warp feature by adopting a Flow Infused Attention module on virtual try-on. The dense warp flow map is projected as indirect guidance attention to enhance the feature map warping in the generation process implicitly, which is less sensitive to the warping estimation accuracy than an explicit warp of the garment image. To further enhance implicit warp guidance, we incorporate high-level spatial attention to complement the dense warp. Experimental results on the VTON-HD and DressCode dataset significantly outperform state-of-the-art methods, demonstrating that FIA-VTON is effective and robust for virtual try-on.

虚拟试穿隐式变形注意力机制

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。