arXiv:2512.20340cs.CV2025-12中稿 · CVPR

通过关键帧注入细节,提升虚拟试穿视频的服装动态与背景一致性。

The devil is in the details: Enhancing Video Virtual Try-On via Keyframe-Driven Details Injection

  • 用关键帧引导采样,提取服装动态和背景一致性信息。
  • 在标准扩散模型中注入增强细节,实现高保真试穿视频生成。
  • 适用于需要高质量服装动画的电商与元宇宙场景。

尽管基于扩散变换器(DiT)的视频虚拟试穿(VVT)已取得显著进展,现有方法仍难以捕捉精细的服装动态并保持视频帧间的背景完整性。此外,引入额外交互模块导致计算开销增加,且现有公开数据集规模有限、质量不高,制约了模型泛化与有效训练。为此,我们提出新框架KeyTailor及大规模高清数据集ViT-HD。KeyTailor的核心思想是关键帧驱动的细节注入:关键帧天然包含前景动态与背景一致性。具体地,采用指令引导的关键帧采样策略筛选输入视频中的信息帧;随后,设计两个定制模块——服装细节增强模块与协同背景优化模块,分别将服装动态提炼至服装相关潜在表示,并优化背景潜在表示的一致性,均以关键帧为指导。这些增强后的细节被注入标准DiT块,结合姿态、掩码与噪声潜在变量,实现高效且逼真的试穿视频合成。该设计无需显式修改DiT结构,同时避免引入额外复杂度。此外,我们的数据集ViT-HD包含15,070个高清视频样本,分辨率为810×1080,涵盖多样化服装。大量实验表明,KeyTailor在动态与静态场景下均优于当前最优基线,在服装保真度与背景完整性方面表现更优。

原文摘要 · Abstract (English)

Although diffusion transformer (DiT)-based video virtual try-on (VVT) has made significant progress in synthesizing realistic videos, existing methods still struggle to capture fine-grained garment dynamics and preserve background integrity across video frames. They also incur high computational costs due to additional interaction modules introduced into DiTs, while the limited scale and quality of existing public datasets also restrict model generalization and effective training. To address these challenges, we propose a novel framework, KeyTailor, along with a large-scale, high-definition dataset, ViT-HD. The core idea of KeyTailor is a keyframe-driven details injection strategy, motivated by the fact that keyframes inherently contain both foreground dynamics and background consistency. Specifically, KeyTailor adopts an instruction-guided keyframe sampling strategy to filter informative frames from the input video. Subsequently,two tailored keyframe-driven modules, the garment details enhancement module and the collaborative background optimization module, are employed to distill garment dynamics into garment-related latents and to optimize the integrity of background latents, both guided by keyframes.These enriched details are then injected into standard DiT blocks together with pose, mask, and noise latents, enabling efficient and realistic try-on video synthesis. This design ensures consistency without explicitly modifying the DiT architecture, while simultaneously avoiding additional complexity. In addition, our dataset ViT-HD comprises 15, 070 high-quality video samples at a resolution of 810*1080, covering diverse garments. Extensive experiments demonstrate that KeyTailor outperforms state-of-the-art baselines in terms of garment fidelity and background integrity across both dynamic and static scenarios.

视频生成虚拟试穿扩散模型关键帧

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。