用语义引导替代传统几何模块,实现无缝视频虚拟试穿。
UniVVT: A Unified End-to-End Framework for High-Fidelity Video Virtual Try-on

- 将视频虚拟试穿重构为语义条件生成,端到端完成
- 在多个基准上达到顶尖性能,优于现有方法
- 适合需要轻量部署与高真实感的电商试穿场景
视频虚拟试穿(VVT)旨在合成一个人穿着目标服装的视频,同时保留身份、动作和场景动态。主流方法将VVT视为掩码条件视频修复,依赖独立的人体分割、姿态估计和服装变形模块。这种多阶段设计增加了部署复杂性,且显式几何先验中的误差会不可逆地传播至生成视频。我们提出UniVVT,一种统一的端到端框架,将VVT重构为语义条件视频生成,推理时完全消除掩码、姿态和变形模块。核心是一个基于多模态大语言模型的场景-任务感知器,联合编码源视频、目标服装和任务指令,生成紧凑的任务感知隐式特征,隐式捕捉转移内容、位置与方式。一个轻量级语义桥将这些特征对齐至基于扩散模型的视频生成器的条件空间,实现连贯的服装迁移。为稳健耦合异构组件,我们设计三阶段渐进训练策略:语义对齐、联合任务适应与灵活分辨率优化。大量实验表明,UniVVT在多个基准上均达到最先进性能,验证了隐式语义引导作为端到端虚拟试穿中脆弱几何预处理的简单有效替代方案。
原文摘要 · Abstract (English)
Video Virtual Try-On (VVT) synthesizes a video of a person wearing a target garment while preserving identity, motion, and scene dynamics. Dominant approaches cast VVT as mask-conditioned video inpainting and rely on separate modules for human parsing, pose estimation, and garment warping. This multi-stage design complicates deployment and, more critically, allows errors in explicit geometric priors to propagate irreversibly into the generated video. We present UniVVT, a unified end-to-end framework that reframes VVT as semantically conditioned video generation, eliminating mask, pose, and warping modules at inference. At its core, a scene-task perceiver built on a Multimodal Large Language Model jointly encodes the source video, target garment, and task instruction into compact, task-aware latent tokens, implicitly capturing what to transfer and where and how to transfer it. A lightweight semantic bridge then aligns these tokens with the conditioning space of a diffusion-based video generator, enabling coherent garment transfer. To robustly couple the heterogeneous components, we devise a three-stage progressive training strategy comprising semantic alignment, joint task adaptation, and flexible-resolution refinement. Extensive experiments demonstrate that UniVVT achieves state-of-the-art performance across multiple benchmarks, validating implicit semantic guidance as a simple and effective alternative to fragile geometric preprocessing for end-to-end virtual try-on.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。