arXiv:2412.09822cs.CV2024-12中稿 · The 36th British M…被引 5

用动态注意力机制提升复杂动作下的虚拟试穿视频稳定性

Dynamic Try-On: Taming Video Virtual Try-on with Dynamic Attention Mechanism

  • 用DiT模型自身充当服装编码器,降低计算开销
  • 动态特征融合模块保留服装细节,生成更稳定结果
  • 肢体感知注意力确保快速动作中身体部位的一致性

视频虚拟试穿具有巨大实际应用潜力,但现有方法在复杂动作下表现不佳。传统方法依赖额外服装编码器,增加计算负担。本文提出基于扩散变换器(Diffusion Transformer, DiT)的Dynamic Try-On框架:利用DiT主干作为服装编码器,并引入动态特征融合模块存储与整合服装特征以降低资源消耗;同时设计肢体感知动态注意力模块,在去噪过程中引导模型关注人体四肢区域,提升复杂运动下的时序一致性。大量实验表明,该方法在复杂姿态视频上仍能生成稳定流畅的试穿效果。

原文摘要 · Abstract (English)

Video try-on stands as a promising area for its tremendous real-world potential. Previous research on video try-on has primarily focused on transferring product clothing images to videos with simple human poses, while performing poorly with complex movements. To better preserve clothing details, those approaches are armed with an additional garment encoder, resulting in higher computational resource consumption. The primary challenges in this domain are twofold: (1) leveraging the garment encoder's capabilities in video try-on while lowering computational requirements; (2) ensuring temporal consistency in the synthesis of human body parts, especially during rapid movements. To tackle these issues, we propose a novel video try-on framework based on Diffusion Transformer(DiT), named Dynamic Try-On. To reduce computational overhead, we adopt a straightforward approach by utilizing the DiT backbone itself as the garment encoder and employing a dynamic feature fusion module to store and integrate garment features. To ensure temporal consistency of human body parts, we introduce a limb-aware dynamic attention module that enforces the DiT backbone to focus on the regions of human limbs during the denoising process. Extensive experiments demonstrate the superiority of Dynamic Try-On in generating stable and smooth try-on results, even for videos featuring complicated human postures.

视频生成扩散模型虚拟试穿

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。