arXiv:2501.08682cs.CVcs.GR2025-01被引 10

让虚拟试衣视频更真实,保持衣服穿在人身上的一致性。

RealVVT: Towards Photorealistic Video Virtual Try-on via Spatio-Temporal Consistency

  • 用时空一致性策略和注意力损失,确保衣服在视频中不扭曲变形。
  • 在多个数据集上超越现有模型,视频试衣效果更自然流畅。
  • 适合电商虚拟试衣、服装设计等需要长期视频一致性的场景。

虚拟试衣是计算机视觉与时尚交叉的关键任务,旨在数字模拟衣物在人体上的穿着效果。尽管单图虚拟试衣已有显著进展,现有方法在长视频序列中难以保持衣物外观的稳定与真实,主要因动态姿态变化和目标衣物特征难以维持。为此,我们利用预训练视频基础模型,提出RealVVT——一个专注于提升动态视频环境中稳定性与真实感的逼真视频虚拟试衣框架。方法包括衣物与时间一致性策略、无感引导注意力聚焦损失机制以保证空间一致性,以及姿态引导的长视频虚拟试衣技术,有效处理长时间视频序列。在多个数据集上的广泛实验表明,该方法在单图与视频虚拟试衣任务中均优于当前最先进模型,为时尚电商与虚拟试衣环境提供了实用解决方案。

原文摘要 · Abstract (English)

Virtual try-on has emerged as a pivotal task at the intersection of computer vision and fashion, aimed at digitally simulating how clothing items fit on the human body. Despite notable progress in single-image virtual try-on (VTO), current methodologies often struggle to preserve a consistent and authentic appearance of clothing across extended video sequences. This challenge arises from the complexities of capturing dynamic human pose and maintaining target clothing characteristics. We leverage pre-existing video foundation models to introduce RealVVT, a photoRealistic Video Virtual Try-on framework tailored to bolster stability and realism within dynamic video contexts. Our methodology encompasses a Clothing & Temporal Consistency strategy, an Agnostic-guided Attention Focus Loss mechanism to ensure spatial consistency, and a Pose-guided Long Video VTO technique adept at handling extended video sequences.Extensive experiments across various datasets confirms that our approach outperforms existing state-of-the-art models in both single-image and video VTO tasks, offering a viable solution for practical applications within the realms of fashion e-commerce and virtual fitting environments.

虚拟试衣视频生成一致性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。