arXiv:2508.02807cs.CV2025-08被引 14

用分阶段扩散模型实现真实场景下逼真的视频虚拟试穿

DreamVVT: Mastering Realistic Video Virtual Try-On in the Wild via a Stage-Wise Diffusion Transformer Framework

  • 分两阶段设计:先生成关键帧,再基于运动与外观描述合成视频
  • 在真实场景下保持服装细节和长时间时序一致性,优于现有方法
  • 适合电商、娱乐领域开发者,尤其关注动态试穿效果的用户

视频虚拟试穿(VVT)技术因在电商广告和娱乐领域的潜力而受到广泛关注。然而,现有端到端方法严重依赖稀缺的成对服装中心数据集,难以有效利用先进视觉模型的先验知识和测试时输入信息,导致在非受限场景中难以准确保留精细服装细节并维持时间一致性。为此,我们提出DreamVVT,一个基于扩散变换器(DiTs)的两阶段框架,可自然利用多样化的非配对人体中心数据以提升真实场景适应性。第一阶段从输入视频采样代表性帧,结合视觉语言模型(VLM)的多帧试穿模型,生成高保真且语义一致的关键帧试穿图像,作为后续视频生成的外观引导。第二阶段从输入内容提取骨架图、细粒度运动与外观描述,连同关键帧试穿图像输入增强式预训练视频生成模型(集成LoRA适配器),确保未见区域的长期时序连贯性,实现高度可信的动态动作。大量定量与定性实验表明,DreamVVT在真实场景下显著优于现有方法,有效保留服装细节并提升时间稳定性。

原文摘要 · Abstract (English)

Video virtual try-on (VVT) technology has garnered considerable academic interest owing to its promising applications in e-commerce advertising and entertainment. However, most existing end-to-end methods rely heavily on scarce paired garment-centric datasets and fail to effectively leverage priors of advanced visual models and test-time inputs, making it challenging to accurately preserve fine-grained garment details and maintain temporal consistency in unconstrained scenarios. To address these challenges, we propose DreamVVT, a carefully designed two-stage framework built upon Diffusion Transformers (DiTs), which is inherently capable of leveraging diverse unpaired human-centric data to enhance adaptability in real-world scenarios. To further leverage prior knowledge from pretrained models and test-time inputs, in the first stage, we sample representative frames from the input video and utilize a multi-frame try-on model integrated with a vision-language model (VLM), to synthesize high-fidelity and semantically consistent keyframe try-on images. These images serve as complementary appearance guidance for subsequent video generation. \textbf{In the second stage}, skeleton maps together with fine-grained motion and appearance descriptions are extracted from the input content, and these along with the keyframe try-on images are then fed into a pretrained video generation model enhanced with LoRA adapters. This ensures long-term temporal coherence for unseen regions and enables highly plausible dynamic motions. Extensive quantitative and qualitative experiments demonstrate that DreamVVT surpasses existing methods in preserving detailed garment content and temporal stability in real-world scenarios. Our project page https://virtu-lab.github.io/

视频生成虚拟试穿扩散模型时序一致性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。