无需额外结构提示,仅靠指令和参考图实现精准视频试穿。
InstructVVT: Instruction-Driven Video Virtual Try-On without Auxiliary Spatial Priors

- 通过双层参考条件机制,从输入三元组中直接提取细粒度控制信号。
- 在ViViD-S和TripVVT-Bench上优于现有开源方法,保持服装还原度与时序一致性。
- 适合追求少干预、高自然度视频试穿的开发者与设计师。
视频虚拟试穿是一项高度受限的编辑任务,需精确替换目标人物的服装,同时严格保留原始视频的空间结构与时间动态。现有方法严重依赖手工设计的空间先验(如掩码、姿态)进行编辑控制,但这些先验在真实场景下易失效,常将丰富视觉上下文压缩为不完整的结构信号。此外,标准重建目标无法充分捕捉试穿特有的人机偏好。为此,我们提出InstructVVT,一种基于扩散Transformer(DiT)的指令驱动、参考引导的视频虚拟试穿框架,无需推理时空间先验。核心思想是通过双级参考条件机制,从输入三元组(源视频、参考服饰、指令)中直接恢复细粒度控制。具体地,一个多模态大模型(MLLM)推断语义编辑令牌以实现目标消歧与结构保留,轻量级条件路径则显式注入细粒度服饰视觉细节。最后,设计试穿专用奖励函数,并采用DiffusionNFT算法对齐人类偏好。在ViViD-S与TripVVT-Bench上的大量实验表明,InstructVVT在服装保真度、结构保留与时序一致性方面均优于当前最先进的开源方法,且所需推理控制更少。
原文摘要 · Abstract (English)
Video virtual try-on is a highly constrained editing task requiring the precise replacement of a target person's clothing while strictly preserving the original video's spatial structure and temporal dynamics. Existing methods heavily rely on auxiliary handcrafted spatial priors (e.g., masks, poses) for editing control. However, these priors are prone to failure in unconstrained real-world videos and often compress rich visual context into incomplete structural signals. Furthermore, standard reconstruction objectives fail to fully capture try-on-specific human preferences. To address these challenges, we propose InstructVVT, an instruction-driven and reference-guided video virtual try-on framework based on a Diffusion Transformer (DiT) that operates without inference-time spatial priors. Our core insight is to recover fine-grained control directly from the input triplet (source video, reference garment, and instruction) via a dual-level reference conditioning scheme. Specifically, an MLLM infers semantic edit tokens for target disambiguation and structural preservation, while a lightweight conditioning pathway explicitly injects fine-grained visual garment details. Finally, we design a try-on-specific reward and utilize the DiffusionNFT algorithm to align the model with human preferences. Extensive experiments on ViViD-S and TripVVT-Bench demonstrate that InstructVVT outperforms state-of-the-art open-source methods in garment fidelity, structural preservation, and temporal consistency, despite requiring fewer inference-time controls.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。