arXiv:2604.27958cs.CV2026-04被引 2

构建首个万级真实场景试穿数据集,提升视频试穿真实性和稳定性。

TripVVT: A Large-Scale Triplet Dataset and a Coarse-Mask Baseline for In-the-Wild Video Virtual Try-On

论文配图:TripVVT: A Large-Scale Triplet Dataset and a Coarse-Mask Baseline for In-the-Wild Video Virtual Try-On
图 1 · 摘自论文原文
  • 用人体掩码先验替代易错的衣物掩码,增强鲁棒性。
  • 在100种复杂场景下,视频质量与试穿保真度显著优于现有方法。
  • 适合做视频试穿、虚拟穿搭系统研究者参考。

由于缺乏大规模真实场景三元组数据及掩码使用不当,视频虚拟试穿模型性能受限。本文首次提出**TripVVT-10K**,目前最大最多样化的真实场景三元组数据集,提供视频级别的跨服装监督,填补现有数据集空白。基于此,我们构建**TripVVT**——一种基于扩散变换器的框架,以简单稳定的人体掩码先验替代脆弱的衣物掩码,有效保留背景,且对真实运动、遮挡和复杂场景具有强鲁棒性。为支持全面评估,我们进一步建立**TripVVT-Bench**,包含100个测试案例,涵盖多样服装、复杂环境及多人场景,评估指标覆盖视频质量、试穿保真度、背景一致性与时间连贯性。相比先进学术与商业系统,TripVVT在视频质量与服装保真度上表现更优,并显著提升对挑战性真实视频的泛化能力。我们公开发布数据集与基准,为可控、逼真、时序稳定的视频虚拟试穿研究奠定坚实基础。

原文摘要 · Abstract (English)

Due to the scarcity of large-scale in-the-wild triplet data and the improper use of masks, the performance of video virtual try-on models remains limited. In this paper, we first introduce **TripVVT-10K**, the largest and most diverse in-the-wild triplet dataset to date, providing explicit video-level cross-garment supervision that existing video datasets lack. Built upon this resource, we develop **TripVVT**, a Diffusion Transformer-based framework that replaces fragile garment masks with a simple, stable human-mask prior, enabling reliable background preservation while remaining robust to real-world motion, occlusion, and cluttered scenes. To support comprehensive evaluation, we further establish **TripVVT-Bench**, a 100-case benchmark covering diverse garments, complex environments, and multi-person scenarios, with metrics spanning video quality, try-on fidelity, background consistency, and temporal coherence. Compared to state-of-the-art academic and commercial systems, TripVVT achieves superior video quality and garment fidelity while markedly improving generalization to challenging in-the-wild videos. We publicly release the dataset and benchmark, which we believe provide a solid foundation for advancing controllable, realistic, and temporally stable video virtual try-on.

视频试穿扩散模型数据集虚拟穿搭

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。