arXiv:2510.07654cs.CV2025-10被引 1

仅用一次换装就实现高效视频虚拟试穿,节省大量参数和计算。

Once Is Enough: Lightweight DiT-Based Video Virtual Try-On via One-Time Garment Appearance Injection

  • 基于首帧换装,利用图像模型替换初始帧服装。
  • 通过姿态与掩码控制生成后续帧,保持时序一致性。
  • 参数与计算量大幅降低,性能仍达领先水平。

视频虚拟试穿旨在将视频中人物的服装替换为目标服饰。现有基于U-Net的双分支扩散模型虽取得显著进展,但将其适配至基于Diffusion Transformer的架构仍具挑战:首先,引入服装参考分支的潜在特征需修改或添加主干网络,导致大量可训练参数;其次,服装潜在特征缺乏固有时序特性,需额外学习。为此,我们提出OIE(Once is Enough)策略,一种基于首帧服装替换的虚拟试穿方法:先使用图像级服装迁移模型替换首帧服装,再以编辑后的首帧内容为控制,结合姿态与掩码信息引导视频生成模型,按序合成剩余帧。实验表明,该方法在保持领先性能的同时,显著提升参数效率与计算效率。

原文摘要 · Abstract (English)

Video virtual try-on aims to replace the clothing of a person in a video with a target garment. Current dual-branch architectures have achieved significant success in diffusion models based on the U-Net; however, adapting them to diffusion models built upon the Diffusion Transformer remains challenging. Initially, introducing latent space features from the garment reference branch requires adding or modifying the backbone network, leading to a large number of trainable parameters. Subsequently, the latent space features of garments lack inherent temporal characteristics and thus require additional learning. To address these challenges, we propose a novel approach, OIE (Once is Enough), a virtual try-on strategy based on first-frame clothing replacement: specifically, we employ an image-based clothing transfer model to replace the clothing in the initial frame, and then, under the content control of the edited first frame, utilize pose and mask information to guide the temporal prior of the video generation model in synthesizing the remaining frames sequentially. Experiments show that our method achieves superior parameter efficiency and computational efficiency while still maintaining leading performance under these constraints.

视频生成虚拟试穿扩散模型轻量化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。