arXiv:2412.03021cs.CVcs.AI2024-12

无需遮罩的视频试穿,用关键点精准控制衣物迁移与时间连贯性。

PEMF-VTO: Point-Enhanced Video Virtual Try-on via Mask-free Paradigm

  • 通过稀疏关键点对齐实现无遮罩的衣物迁移引导。
  • 在野外复杂场景下生成更自然、连贯的试穿视频,显著提升视觉质量。
  • 适合需要高精度动态人体试穿的电商与虚拟试衣应用。

视频虚拟试穿旨在将参考服装无缝地转移到视频中目标人物身上,同时保持视觉保真度和时间连贯性。现有方法通常依赖修复遮罩定义试穿区域,适用于简单场景(如店内视频)。然而,在复杂真实场景中,过大的不一致遮罩会破坏时空信息,导致结果失真。无遮罩方法虽缓解此问题,但在动态身体动作下难以准确确定试穿区域。为此,我们提出 PEMF-VTO,一种基于稀疏点对齐的新型无遮罩视频虚拟试穿框架。其核心创新在于引入点增强引导,灵活可靠地控制空间级衣物迁移与时间级视频连贯性。具体设计了点增强变压器(PET),包含两个组件:点增强空间注意力(PSA),利用帧-衣关键点对齐精确引导衣物迁移;点增强时间注意力(PTA),利用帧-帧关键点对应关系增强时间连贯性,确保帧间过渡平滑。大量实验表明,PEMF-VTO优于现有最先进方法,在挑战性野外场景中生成更自然、连贯且视觉吸引力强的试穿视频。

原文摘要 · Abstract (English)

Video Virtual Try-on aims to seamlessly transfer a reference garment onto a target person in a video while preserving both visual fidelity and temporal coherence. Existing methods typically rely on inpainting masks to define the try-on area, enabling accurate garment transfer for simple scenes (e.g., in-shop videos). However, these mask-based approaches struggle with complex real-world scenarios, as overly large and inconsistent masks often destroy spatial-temporal information, leading to distorted results. Mask-free methods alleviate this issue but face challenges in accurately determining the try-on area, especially for videos with dynamic body movements. To address these limitations, we propose PEMF-VTO, a novel Point-Enhanced Mask-Free Video Virtual Try-On framework that leverages sparse point alignments to explicitly guide garment transfer. Our key innovation is the introduction of point-enhanced guidance, which provides flexible and reliable control over both spatial-level garment transfer and temporal-level video coherence. Specifically, we design a Point-Enhanced Transformer (PET) with two core components: Point-Enhanced Spatial Attention (PSA), which uses frame-cloth point alignments to precisely guide garment transfer, and Point-Enhanced Temporal Attention (PTA), which leverages frame-frame point correspondences to enhance temporal coherence and ensure smooth transitions across frames. Extensive experiments demonstrate that our PEMF-VTO outperforms state-of-the-art methods, generating more natural, coherent, and visually appealing try-on videos, particularly for challenging in-the-wild scenarios. The link to our paper's homepage is https://pemf-vto.github.io/.

视频试穿无遮罩关键点引导时空连贯

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。