将衣架视频转为逼真人像,保持身份一致且动作自然。
From Mannequin to Human: A Pose-Aware and Identity-Preserving Video Generation Framework for Lifelike Clothing Display
- 用姿态感知编码器融合面部与姿势信息,确保身份稳定。
- 通过像素空间镜像损失修复细节,提升面部清晰度。
- 适合虚拟试衣、电商展示等需要真实感的场景。
衣架式服装展示成本低,但缺乏真实感和表现力。为此,我们提出新任务——衣架到真人(M2H)视频生成,旨在从衣架视频合成可控制身份的逼真人像视频。提出M2HVideo框架,解决头部与身体动作不匹配、时序建模导致的身份漂移两大挑战。引入动态姿态感知头编码器,融合面部语义与身体姿态,生成跨帧一致的身份嵌入;为缓解潜在空间压缩造成的面部细节丢失,设计基于DDIM的一步去噪像素空间镜像损失;并构建分布感知适配器,对齐身份与服饰特征的统计分布,增强时序一致性。在UBC时尚数据集、自建ASOS数据集及现场采集的MannequinVideos数据集上的实验表明,M2HVideo在服饰一致性、身份保留与视频保真度方面均优于现有方法。
原文摘要 · Abstract (English)
Mannequin-based clothing displays offer a cost-effective alternative to real-model showcases for online fashion presentation, but lack realism and expressive detail. To overcome this limitation, we introduce a new task called mannequin-to-human (M2H) video generation, which aims to synthesize identity-controllable, photorealistic human videos from footage of mannequins. We propose M2HVideo, a pose-aware and identity-preserving video generation framework that addresses two key challenges: the misalignment between head and body motion, and identity drift caused by temporal modeling. In particular, M2HVideo incorporates a dynamic pose-aware head encoder that fuses facial semantics with body pose to produce consistent identity embeddings across frames. To address the loss of fine facial details due to latent space compression, we introduce a mirror loss applied in pixel space through a denoising diffusion implicit model (DDIM)-based one-step denoising. Additionally, we design a distribution-aware adapter that aligns statistical distributions of identity and clothing features to enhance temporal coherence. Extensive experiments on the UBC fashion dataset, our self-constructed ASOS dataset, and the newly collected MannequinVideos dataset captured on-site demonstrate that M2HVideo achieves superior performance in terms of clothing consistency, identity preservation, and video fidelity in comparison to state-of-the-art methods.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。