arXiv:2505.21325cs.CV2025-05被引 11

用扩散模型提升试穿视频的细节和时序一致性

MagicTryOn: Harnessing Diffusion Transformer for Garment-Preserving Video Virtual Try-on

  • 分解服装特征注入去噪过程,保留细粒度纹理
  • 引入时空旋转位置编码,减少帧间抖动与外观漂移
  • 支持实时推理,适合动态场景下的虚拟试穿应用

视频虚拟试穿(VVT)旨在生成连续帧中自然呈现的服装图像,同时捕捉其动态变化与人体动作的交互。现有方法仍存在服装保真度不足和时空一致性差的问题。原因在于:(1) 服装信息利用不充分,仅注入有限的服装提示,导致细节保真度弱;(2) 缺乏有效的时空建模,影响跨帧身份一致性和时间稳定性,引发抖动与外观漂移。本文提出 MagicTryOn,一种基于扩散变换器的服装保真视频虚拟试穿框架。为保留精细服装细节,设计细粒度服装保真策略,解耦服装提示并注入去噪过程。为提升时序一致性、抑制抖动,提出服装感知的时空旋转位置编码(RoPE),在全自注意力中扩展 RoPE,利用时空相对位置调制服装令牌。训练时引入掩码感知损失,增强服装区域的保真度。此外,采用分布匹配蒸馏将采样轨迹压缩至四步,实现实时推理且不降低服装保真度。大量定量与定性实验表明,MagicTryOn 在非受限场景下优于现有方法,显著提升服装细节保真度与时序稳定性。

原文摘要 · Abstract (English)

Video Virtual Try-On (VVT) aims to synthesize garments that appear natural across consecutive video frames, capturing both their dynamics and interactions with human motion. Despite recent progress, existing VVT methods still suffer from inadequate garment fidelity and limited spatiotemporal consistency. The reasons are: (1) under-exploitation of garment information, with limited garment cues being injected, resulting in weaker fine-detail fidelity; and (2) a lack of spatiotemporal modeling, which hampers cross-frame identity consistency and causes temporal jitter and appearance drift. In this paper, we present MagicTryOn, a diffusion-transformer based framework for garment-preserving video virtual try-on. To preserve fine-grained garment details, we propose a fine-grained garment-preservation strategy that disentangles garment cues and injects these decomposed priors into the denoising process. To improve temporal garment consistency and suppress jitter, we introduce a garment-aware spatiotemporal rotary positional embedding (RoPE) that extends RoPE within full self-attention, using spatiotemporal relative positions to modulate garment tokens. We further impose a mask-aware loss during training to enhance fidelity within garment regions. Moreover, we adopt distribution-matching distillation to compress the sampling trajectory to four steps, enabling real-time inference without degrading garment fidelity. Extensive quantitative and qualitative experiments demonstrate that MagicTryOn outperforms existing methods, delivering superior garment-detail fidelity and temporal stability in unconstrained settings.

虚拟试穿扩散模型视频生成服装保真

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。