实时高保真虚拟试衣,速度提升26倍且保持画质。
LiveVVT: High-Fidelity Video Virtual Try-On in Real Time

- 滚动窗口内有限前瞻双向建模,兼顾时序一致性与实时性。
- 引入短期动态记忆与长期外观记忆,确保服装细节稳定。
- 通过渐进式蒸馏优化推理流程,适配连续生成场景。
基于扩散模型的视频虚拟试衣(VVT)通过双向时空建模实现高视觉保真度,但全片段依赖导致实际连续部署中延迟和计算开销过高。盲目施加因果性会破坏预训练的双向先验,显著降低生成质量。本文提出LiveVVT,一种滚动流式扩散框架,在因果递归生成中保留受限的双向建模能力。在固定大小窗口内,LiveVVT联合去噪多个视频片段,支持有限前瞻,保持局部双向交互并每轮输出一个干净片段。窗口外,两种互补记忆维持长期一致性:有限时序记忆传播近期动态与遮挡上下文,而持久全局外观记忆仅需一次构建,来自目标服饰和正脸试穿关键帧,锚定服装细节与整体着装外观。此外,提出渐进式蒸馏框架,融合双向VVT学习、教师轨迹回归以实现因果少步适应,以及协同匹配蒸馏,将教师分布匹配与真实视频上的滚动流匹配结合,使优化对齐递归推理。在配对与非配对长序列基准上实验表明,相比同规模模型,生成质量更优,延迟降低26倍,吞吐量提高11倍,实现高保真实时流式VVT。
原文摘要 · Abstract (English)
Diffusion-based Video Virtual Try-On (VVT) achieves high visual fidelity through bidirectional spatio-temporal modeling, but complete-clip dependence incurs prohibitive latency and computational overhead in practical continuous deployment. Naively enforcing causality disrupts pretrained bidirectional priors and substantially degrades synthesis quality. We introduce LiveVVT, a rolling streaming diffusion framework that preserves bounded bidirectional modeling within causal recurrent generation. Within a fixed-size window, LiveVVT jointly denoises multiple video chunks under bounded look-ahead, preserving local bidirectional interactions while emitting one clean chunk per iteration. Beyond the window, two complementary memories sustain long-term consistency: a bounded temporal memory propagates recent dynamics and occlusion context, whereas a persistent global appearance memory, constructed once from the target garment and a frontal try-on keyframe, anchors garment details and dressed appearance throughout the stream. We further introduce a progressive distillation framework integrating bidirectional VVT learning, teacher-trajectory regression for causal few-step adaptation, and Collaborative Matching Distillation, which couples teacher-distribution matching with rolling flow matching on real videos to align optimization with recurrent inference. Experiments on paired and unpaired long-sequence benchmarks demonstrate superior generation quality over similarly sized models, with $26\times$ lower latency and $11\times$ higher throughput, enabling high-fidelity real-time streaming VVT.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。