arXiv:2504.17414cs.CV2025-04被引 3

用3D纹理网格引导扩散模型,实现服装试穿的高清连贯效果

3DV-TON: Textured 3D-Guided Consistent Video Try-on via Diffusion Models

  • 通过动态3D纹理网格提供逐帧一致的服装运动指导
  • 在高分辨率数据集HR-VVT上达到优于现有方法的生成质量
  • 适合关注视频服装替换与运动一致性的研究者

视频试穿将视频中的衣物更换为目标服装。现有方法在处理复杂图案和多样人体姿态时,难以生成高质量且时间连贯的结果。我们提出3DV-TON,一种基于扩散模型的新框架,可生成高保真且时间一致的视频试穿结果。该方法利用可驱动的纹理3D网格作为显式逐帧引导,缓解模型过度关注外观保真度而忽视动作连贯性的问题,通过直接参考整个视频序列中一致的服装纹理运动来实现。所提方法包含自适应生成动态3D引导的流程:(1) 选择关键帧进行初始2D图像试穿,(2) 重建并同步原始视频姿态的动画纹理3D网格。此外,引入稳健的矩形掩码策略,有效抑制因人体与服装动态运动导致的服装信息泄漏引发的伪影传播。为推动视频试穿研究,我们构建了高分辨率基准数据集HR-VVT,包含130个视频,涵盖多种服装类型与场景。定量与定性结果表明,本方法显著优于现有方法。

原文摘要 · Abstract (English)

Video try-on replaces clothing in videos with target garments. Existing methods struggle to generate high-quality and temporally consistent results when handling complex clothing patterns and diverse body poses. We present 3DV-TON, a novel diffusion-based framework for generating high-fidelity and temporally consistent video try-on results. Our approach employs generated animatable textured 3D meshes as explicit frame-level guidance, alleviating the issue of models over-focusing on appearance fidelity at the expanse of motion coherence. This is achieved by enabling direct reference to consistent garment texture movements throughout video sequences. The proposed method features an adaptive pipeline for generating dynamic 3D guidance: (1) selecting a keyframe for initial 2D image try-on, followed by (2) reconstructing and animating a textured 3D mesh synchronized with original video poses. We further introduce a robust rectangular masking strategy that successfully mitigates artifact propagation caused by leaking clothing information during dynamic human and garment movements. To advance video try-on research, we introduce HR-VVT, a high-resolution benchmark dataset containing 130 videos with diverse clothing types and scenarios. Quantitative and qualitative results demonstrate our superior performance over existing methods. The project page is at this link https://2y7c3.github.io/3DV-TON/

视频试穿扩散模型3D引导运动一致性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。