用注意力引导生成更连贯且细节丰富的虚拟试衣视频。
ChronoTailor: Harnessing Attention Guidance for Fine-Grained Video Virtual Try-On
- 通过时空注意力机制精准融合服装细节特征。
- 在动态动作中保持服装形状与视频连贯性,减少伪影。
- 适合服装设计、虚拟试衣系统研发人员使用。
视频虚拟试衣旨在无缝将源视频中人物的衣物替换为目标服装。尽管该领域已有显著进展,现有方法仍难以保持时序连续性并还原服装细节。本文提出基于扩散模型的ChronoTailor框架,在保证时序一致性的同时精确保留服装细节。首先,ChronoTailor采用区域感知的空间引导机制调控空间注意力演化,并引入注意力驱动的时序特征融合机制,生成更连贯的时序特征,既支持局部精细编辑,又有效抑制视频动态带来的伪影。其次,通过多尺度服装特征融合保留低层视觉细节,并引入服装-姿态特征对齐策略确保动态运动中的时序连续性。此外,我们构建了新数据集StyleDress,包含复杂服饰、多样环境与多种姿势,优于现有公开数据集,将向学术界开放。大量实验表明,ChronoTailor在动作过程中显著提升时序连续性与服装细节保真度,明显优于先前方法。
原文摘要 · Abstract (English)
Video virtual try-on aims to seamlessly replace the clothing of a person in a source video with a target garment. Despite significant progress in this field, existing approaches still struggle to maintain continuity and reproduce garment details. In this paper, we introduce ChronoTailor, a diffusion-based framework that generates temporally consistent videos while preserving fine-grained garment details. By employing a precise spatio-temporal attention mechanism to guide the integration of fine-grained garment features, ChronoTailor achieves robust try-on performance. First, ChronoTailor leverages region-aware spatial guidance to steer the evolution of spatial attention and employs an attention-driven temporal feature fusion mechanism to generate more continuous temporal features. This dual approach not only enables fine-grained local editing but also effectively mitigates artifacts arising from video dynamics. Second, ChronoTailor integrates multi-scale garment features to preserve low-level visual details and incorporates a garment-pose feature alignment to ensure temporal continuity during dynamic motion. Additionally, we collect StyleDress, a new dataset featuring intricate garments, varied environments, and diverse poses, offering advantages over existing public datasets, and will be publicly available for research. Extensive experiments show that ChronoTailor maintains spatio-temporal continuity and preserves garment details during motion, significantly outperforming previous methods.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。