用单图生成任意时长的虚拟试穿视频,支持自然动作与连贯视觉。
Virtual Fitting Room: Generating Arbitrarily Long Videos of Virtual Try-On from a Single Image -- Technical Preview
- 分段自回归生成视频,无需长视频数据
- 生成分钟级视频,局部平滑且全局一致
- 适合虚拟试衣、数字人展示等场景
我们提出虚拟试穿间(Virtual Fitting Room, VFR),一种新颖的视频生成模型,可生成任意长度的虚拟试穿视频。将长视频生成任务建模为自回归式分段生成过程,避免对资源密集型生成和长视频数据的需求,同时具备生成任意长度视频的灵活性。该任务的关键挑战在于:确保相邻片段间的局部平滑性,以及不同片段间的全局时间一致性。为此,我们提出VFR框架,通过前缀视频条件保证平滑性,并利用锚定视频(anchor video)——一个360度全景视频,全面捕捉人体整体外观——来强制时间一致性。在多种动作下,VFR能生成具有局部平滑性和全局时间一致性的分钟级虚拟试穿视频,是长时虚拟试穿视频生成的开创性工作。
原文摘要 · Abstract (English)
We introduce the Virtual Fitting Room (VFR), a novel video generative model that produces arbitrarily long virtual try-on videos. Our VFR models long video generation tasks as an auto-regressive, segment-by-segment generation process, eliminating the need for resource-intensive generation and lengthy video data, while providing the flexibility to generate videos of arbitrary length. The key challenges of this task are twofold: ensuring local smoothness between adjacent segments and maintaining global temporal consistency across different segments. To address these challenges, we propose our VFR framework, which ensures smoothness through a prefix video condition and enforces consistency with the anchor video -- a 360-degree video that comprehensively captures the human's wholebody appearance. Our VFR generates minute-scale virtual try-on videos with both local smoothness and global temporal consistency under various motions, making it a pioneering work in long virtual try-on video generation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。