arXiv:2606.08514cs.CV2026-06

一次生成,任意换装,无需掩码,视频试穿新范式

OmniTryOn: Video Try-On Anything at Once!

论文配图:OmniTryOn: Video Try-On Anything at Once!
图 1 · 摘自论文原文
  • 首帧直接缓存多种服饰,免用外部掩码实现多件同时试穿
  • 通过时空一致位置编码,精准保持人体动作与背景动态
  • 支持复杂场景下多物体同步试穿,适合服装设计与虚拟试衣应用

尽管视频虚拟试穿(VVT)已取得显著进展,现有方法仍存在两大局限:一是仅限单件衣物转移,难以实现多物品同步试穿;二是严重依赖外部先验(如衣物掩码),破坏物理动态并降低视觉质量。为此,本文提出全新的「任意试穿」任务,旨在单次推理中同时将多种可穿戴物转移到视频中的人物上。为支撑该范式,我们构建了包含配对视频数据集和定制评估协议的TryAny-Bench基准。进一步提出OmniTryOn——一种无外部先验的生成框架。该模型采用首帧服饰缓存策略,从初始帧直接获取多样服饰信息;为保证一致性,引入时空一致位置编码(STC-RoPE),建立强鲁棒性时空锚点以严格保留复杂人体运动与背景动态;通过渐进式试穿(GTO)训练策略优化,模型逐步掌握多对象合成能力。在TryAny-Bench上的大量实验表明,OmniTryOn显著优于现有专业视频试穿模型及通用视频编辑基线,确立了「任意试穿」任务的新标准。数据集、代码与模型已开源。

原文摘要 · Abstract (English)

Although video virtual try-on (VVT) has achieved significant progress, existing methods still exhibit two fundamental limitations: first, they are restricted to single-garment transfer, rendering simultaneous multi-object try-on highly impractical; second, their heavy reliance on explicit external priors (e.g., garment masks) inevitably destroys crucial physical dynamics and degrades visual quality. To bridge this gap, this paper proposes the novel Try-On Anything task, which aims to simultaneously transfer diverse wearable objects onto a person in a video in a single inference pass. To support and standardize this paradigm, we introduce TryAny-Bench, a comprehensive benchmark encompassing a paired video dataset alongside a tailored evaluation protocol. Furthermore, we present OmniTryOn, an external-prior-free generative framework designed to tackle this task. Specifically, OmniTryOn employs a First Frame Wearable Cache strategy, which directly provides diverse wearable objects for the generation process through the initial video frame. To maintain consistency, we propose the Spatiotemporally Consistent RoPE (STC-RoPE), which inherently establishes robust spatiotemporal anchors to strictly preserve complex human motions and background dynamics. Optimized by the proposed Gradual Try-On (GTO) training strategy, our model progressively masters robust multi-object synthesis. Extensive experiments on TryAny-Bench demonstrate that OmniTryOn significantly outperforms existing specialized video virtual try-on models and general video editing baselines, establishing a powerful new standard for the Try-On Anything task. Our dataset, code, and models are available at https://github.com/xcltql666/OminTryOn.

视频试穿多物迁移无先验生成模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。