让虚拟试衣更真实:支持人与衣物互动的视频试穿技术
iTryOn: Mastering Interactive Video Virtual Try-On with Spatial-Semantic Guidance

- 通过空间-语义双层引导,精准捕捉手与衣物接触细节
- 在交互式视频试穿任务上超越现有方法,保持动作自然连贯
- 适合想做智能穿搭、虚拟时装展示的研究者和开发者
视频虚拟试穿(VVT)旨在无缝替换视频中人物身上的服装。尽管现有方法在保持时间一致性方面取得进展,但大多局限于非交互场景,仅展示服装效果。本文提出并定义了一个新挑战:交互式视频虚拟试穿(Interactive VVT),即视频中人物主动与衣物互动。该任务带来双重挑战:一是仅靠标准姿态信息难以解析互动的语义模糊性;二是从视频中学习复杂服装形变,而互动时刻稀疏短暂。为此,我们提出iTryOn,基于大规模视频扩散Transformer框架,创新设计多层级交互注入机制。空间层面引入无特定服装的3D手部先验,精确指导手-衣接触;语义层面结合全局描述与时间标记的动作描述,通过新颖的动作用旋转位置编码(A-RoPE)实现同步。大量实验表明,iTryOn不仅在传统VVT基准上达到顶尖性能,更在新交互设置中显著领先,标志着向更动态可控的虚拟试穿体验迈出关键一步。
原文摘要 · Abstract (English)
Video Virtual Try-On (VVT) aims to seamlessly replace a garment on a person in a video with a new one. While existing methods have made significant strides in maintaining temporal consistency, they are predominantly confined to non-interactive scenarios where models merely showcase garments. This limitation overlooks a crucial aspect of real-world apparel presentation: active human-garment interaction. To bridge this gap, we introduce and formalize a new challenging task: Interactive Video Virtual Try-On (Interactive VVT), where subjects in the video actively engage with their clothing. This task introduces unique challenges beyond simple texture preservation, including: (1) resolving the semantic ambiguity of interactions from standard pose information, and (2) learning complex garment deformations from video where interactive moments are sparse and brief. To address these challenges, we propose iTryOn, a novel framework built upon a large-scale video diffusion Transformer. iTryOn pioneers a multi-level interaction injection mechanism to guide the generation of complex dynamics. At the spatial level, we introduce a garment-agnostic 3D hand prior to provide fine-grained guidance for precise hand-garment contact, effectively resolving spatial ambiguity. At the semantic level, iTryOn leverages global captions for overall context and time-stamped action captions for localized interactions, synchronized via our novel Action-aware Rotational Position Embedding (A-RoPE). Extensive experiments demonstrate that iTryOn not only achieves state-of-the-art performance on traditional VVT benchmarks but also establishes a commanding lead in the new interactive setting, marking a significant step towards more dynamic and controllable virtual try-on experiences.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。