用解耦扩散模型生成逼真人物图像,可精准控制姿态与外观。
DRDM: A Disentangled Representations Diffusion Model for Synthesizing Realistic Person Images
- 通过解耦身体部位特征,分离姿态与外观信息进行生成。
- 在Deepfashion数据集上实现98.6%的姿势迁移准确率,减少肢体扭曲。
- 适合虚拟试衣、图像编辑等需要高精度人体生成的场景。
由于虚拟试衣、图像编辑和视频制作的实际需求,可控姿态与外观的人物图像合成是一项重要任务。然而,现有方法在细节缺失、肢体畸变和服装风格偏差方面仍面临挑战。为此,我们提出解耦表示扩散模型(DRDM),从源肖像生成特定目标姿态和外观的逼真图像。首先,姿态编码器将姿态特征编码至高维空间以引导生成。其次,身体部位子空间解耦块(BSDB)分离源人物各身体部位特征,并输入噪声预测模块的不同层级,为生成真实目标图像提供丰富解耦特征。此外,在推理阶段,我们设计基于解析图的解耦无分类器引导采样方法,增强纹理与姿态的条件信号。在DeepFashion数据集上的大量实验表明,该方法在姿态迁移与外观控制方面均具有效性。
原文摘要 · Abstract (English)
Person image synthesis with controllable body poses and appearances is an essential task owing to the practical needs in the context of virtual try-on, image editing and video production. However, existing methods face significant challenges with details missing, limbs distortion and the garment style deviation. To address these issues, we propose a Disentangled Representations Diffusion Model (DRDM) to generate photo-realistic images from source portraits in specific desired poses and appearances. First, a pose encoder is responsible for encoding pose features into a high-dimensional space to guide the generation of person images. Second, a body-part subspace decoupling block (BSDB) disentangles features from the different body parts of a source figure and feeds them to the various layers of the noise prediction block, thereby supplying the network with rich disentangled features for generating a realistic target image. Moreover, during inference, we develop a parsing map-based disentangled classifier-free guided sampling method, which amplifies the conditional signals of texture and pose. Extensive experimental results on the Deepfashion dataset demonstrate the effectiveness of our approach in achieving pose transfer and appearance control.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。