一键生成换装动画视频,避免身份漂移和服装扭曲。
Vanast: Virtual Try-On with Human Image Animation via Synthetic Triplet Supervision

- 单步统一生成换装动画,避免两阶段流程的误差累积。
- 构建大规模合成三元组数据,支持全身服饰与姿态协同建模。
- 零样本换装插值能力,适合电商试穿与虚拟人创作场景。
我们提出 Vanast,一个统一框架,可直接从单张人物图像、服装图像和姿态引导视频生成换装人体动画视频。传统两阶段方法将图像级虚拟试穿与姿态驱动动画分开处理,常导致身份漂移、服装失真及前后不一致问题。本模型通过单步统一过程实现连贯合成。为支持此设定,我们构建大规模合成三元组监督数据:生成保持身份特征但服饰不同于服装目录的图像,捕捉完整的上下装三元组以突破单件服装姿态视频对的限制,并组装无需服装目录图像的多样化真实场景三元组。我们进一步设计双模块架构用于视频扩散变换器,稳定训练,保留预训练生成质量,同时提升服装准确性、姿态遵循度与身份一致性,支持零样本服装插值。这些贡献使 Vanast 能在多种服装类型下生成高保真、身份一致的动画视频。
原文摘要 · Abstract (English)
We present Vanast, a unified framework that generates garment-transferred human animation videos directly from a single human image, garment images, and a pose guidance video. Conventional two-stage pipelines treat image-based virtual try-on and pose-driven animation as separate processes, which often results in identity drift, garment distortion, and front-back inconsistency. Our model addresses these issues by performing the entire process in a single unified step to achieve coherent synthesis. To enable this setting, we construct large-scale triplet supervision. Our data generation pipeline includes generating identity-preserving human images in alternative outfits that differ from garment catalog images, capturing full upper and lower garment triplets to overcome the single-garment-posed video pair limitation, and assembling diverse in-the-wild triplets without requiring garment catalog images. We further introduce a Dual Module architecture for video diffusion transformers to stabilize training, preserve pretrained generative quality, and improve garment accuracy, pose adherence, and identity preservation while supporting zero-shot garment interpolation. Together, these contributions allow Vanast to produce high-fidelity, identity-consistent animation across a wide range of garment types.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。