无需遮罩的视频试穿,用图像伪数据提升真实感与一致性
BooM-VVT: Boosting Mask-Free Video Virtual Try-On with Image-Level Pseudo Data

- 用图像级伪数据替代视频级伪数据,降低训练成本
- 通过服装敏感关键帧采样,保持衣物外观连续性
- 构建多视角数据集,支持复杂视角下的试穿生成
视频虚拟试穿(VVT)旨在生成人物穿戴目标服饰的逼真视频。现有方法虽采用关键帧驱动范式提升野外表现,但仍依赖遮罩定位试穿区域,对大幅运动和严重遮挡敏感。尽管无遮罩图像试穿方法借助大规模伪数据取得进展,但将该范式扩展至视频仍面临挑战——构建视频级伪数据成本过高。此外,粗略的关键帧采样和多视角试穿数据稀缺限制了现有方法在保持衣物一致性与应对多样化任务上的表现。为此,我们提出无遮罩的BooM-VVT框架,基于关键帧驱动范式。为实现无遮罩试穿,引入多阶段训练策略,利用图像级伪数据学习无遮罩定位,显著减少对昂贵视频级伪数据的需求。为提升衣物一致性,提出服装敏感关键帧采样,依据与衣物相关的身体区域选择关键帧以更好捕捉衣物外观。进一步引入帧共享3D-RoPE,建立关键帧与目标视频帧间的时空对应关系,实现精准的衣物细节迁移。最后,构建大型多视角试穿数据集OmniView,支持复杂相机视角下可靠试穿视频生成。大量实验表明,BooM-VVT在时间一致性与衣物保真度上优于现有方法。
原文摘要 · Abstract (English)
Video virtual try-on (VVT) aims to generate realistic videos of a person wearing a target garment. Recent methods leverage a keyframe-driven video generation paradigm to improve in-the-wild performance, yet they still rely on masks to localize try-on regions, making them vulnerable to large motions and severe occlusions. Although mask-free image-based try-on methods have shown promising results by leveraging large-scale pseudo data, extending this paradigm to videos remains difficult, as constructing video-level pseudo data is prohibitively expensive. Furthermore, coarse keyframe sampling and the scarcity of multi-view try-on data limit existing keyframe-driven methods in maintaining garment consistency and handling diverse try-on tasks. To address these challenges, we propose BooM-VVT, a mask-free VVT framework built upon the keyframe-driven paradigm. To achieve mask-free VVT, we introduce a multi-stage training strategy that leverages image-level pseudo data for mask-free localization learning, substantially reducing the need for costly video-level pseudo data. To improve garment consistency, we propose Garment-Sensitive Keyframe Sampling, which selects keyframes based on garment-relevant body regions to better capture garment appearance. We further introduce Frame-Shared 3D-RoPE to establish spatiotemporal correspondences between keyframes and target video frames for accurate garment-detail transfer. Finally, we construct OmniView, a large-scale multi-view try-on dataset to support reliable try-on video generation under complex camera viewpoints and diverse try-on tasks. Extensive experiments demonstrate that BooM-VVT achieves superior temporal consistency and garment fidelity over existing methods. Project page: https://boomvvt.github.io/boomvvt.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。