用动态姿态交互提升视频虚拟试衣的时序一致性
Pursuing Temporal-Consistent Video Virtual Try-On via Dynamic Pose Interaction
- 引入骨架对齐适配器与分层注意力机制,建模人体与服装的时空互动
- 在VVT数据集上实现0.506的VFID分数,较当前最优方法提升60.5%
- 适合关注视频生成中姿态连续性与服装真实感的研究者
视频虚拟试衣旨在将特定服装无缝地套在视频中的人物身上。主要挑战在于保持服装视觉真实性的同时,动态适应人物的姿态与体型变化。现有方法多聚焦于图像级虚拟试衣,直接扩展至视频常导致时序不一致。尽管当前视频试衣方法引入时序模块缓解此问题,仍忽略人体与服装之间关键的时空姿态交互。有效的视频姿态交互不仅需考虑每帧中人体与服装姿态的空间对齐,还需捕捉整个视频中人体姿态的时序动态。为此,本文提出动态姿态交互扩散模型(DPIDM),利用扩散模型深入建模动态姿态交互。技术上,DPIDM引入基于骨架的姿态适配器,将同步的人体与服装姿态融入去噪网络;设计分层注意力模块,通过姿态感知的空间与时间注意力机制,建模帧内人体-服装姿态交互及跨帧长期姿态动态;此外,引入连续帧间的时序正则化注意力损失,增强时序一致性。在VITON-HD、VVT和ViViD数据集上的大量实验表明,所提方法优于基线方法。尤其在VVT数据集上,取得0.506的VFID分数,较当前最优方法GPD-VVTO提升60.5%。
原文摘要 · Abstract (English)
Video virtual try-on aims to seamlessly dress a subject in a video with a specific garment. The primary challenge involves preserving the visual authenticity of the garment while dynamically adapting to the pose and physique of the subject. While existing methods have predominantly focused on image-based virtual try-on, extending these techniques directly to videos often results in temporal inconsistencies. Most current video virtual try-on approaches alleviate this challenge by incorporating temporal modules, yet still overlook the critical spatiotemporal pose interactions between human and garment. Effective pose interactions in videos should not only consider spatial alignment between human and garment poses in each frame but also account for the temporal dynamics of human poses throughout the entire video. With such motivation, we propose a new framework, namely Dynamic Pose Interaction Diffusion Models (DPIDM), to leverage diffusion models to delve into dynamic pose interactions for video virtual try-on. Technically, DPIDM introduces a skeleton-based pose adapter to integrate synchronized human and garment poses into the denoising network. A hierarchical attention module is then exquisitely designed to model intra-frame human-garment pose interactions and long-term human pose dynamics across frames through pose-aware spatial and temporal attention mechanisms. Moreover, DPIDM capitalizes on a temporal regularized attention loss between consecutive frames to enhance temporal consistency. Extensive experiments conducted on VITON-HD, VVT and ViViD datasets demonstrate the superiority of our DPIDM against the baseline methods. Notably, DPIDM achieves VFID score of 0.506 on VVT dataset, leading to 60.5% improvement over the state-of-the-art GPD-VVTO approach.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。