用光流指导视频试穿,无需分割掩码也能保持时序一致
FlowVVTON: Flow-Guided Mask-Free Video Virtual Try-On

- 用光流作为训练监督,跨层对齐相邻帧特征
- 在TikTokDress上时序一致性提升5.7倍(VFID-R)
- 适合做无标注视频试穿,尤其抗大运动和遮挡
视频虚拟试穿旨在将目标服装跨视频帧转移到运动人体上。现有方法依赖人体解析掩码或姿态关键点,在大幅运动和遮挡下常失效,导致边界伪影和时序不一致。此外,多数方法仅用注意力机制建模时序,缺乏显式运动监督。我们提出FlowVVTON,一种完全无掩码的框架,彻底消除对解析掩码的依赖。仅在训练阶段使用光流作为监督信号:跨生成模型所有层施加光流扭曲潜在损失,通过显式物理运动约束对齐相邻帧特征,实现多尺度时序一致性。采用两阶段训练策略,先建立无掩码空间对齐,再引入光流引导的时序监督。在TikTokDress数据集上的实验表明,FlowVVTON显著优于基线方法,尤其在时序一致性上(相比SwiftTry提升5.7× VFID-R),且全程无需分割掩码、姿态关键点或区域标注。
原文摘要 · Abstract (English)
Video virtual try-on aims to transfer a target garment onto a moving person across video frames. Current methods rely on human parsing masks or pose keypoints that frequently fail under large motions and occlusions, causing boundary artifacts and temporal inconsistency. A further limitation is that most approaches rely solely on attention mechanisms for temporal modeling, providing no explicit motion supervision. We propose FlowVVTON, a mask-free framework that eliminates parsing mask dependency entirely. Optical flow is used solely as a training-time supervision signal: a flow-warped latent loss, applied across all layers of the generation model, enforces multi-scale temporal consistency by aligning adjacent-frame features under explicit physical motion constraints. A two-stage training strategy establishes mask-free spatial alignment before introducing flow-guided temporal supervision. Experiments on TikTokDress show that FlowVVTON outperforms baselines by substantial margins, particularly in temporal consistency (5.7$\times$ VFID-R improvement over SwiftTry), while requiring no segmentation masks, pose keypoints, or region annotations at any stage.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。