通过显式对齐融合嵌入,提升姿势引导人像生成的细节保真度与一致性
Fusion Embedding for Pose-Guided Person Image Synthesis with Diffusion Model
- 设计显式融合嵌入机制,用对比学习对齐源图与姿态特征
- 在DeepFashion和RWTH-PHOENIX数据集上实现更优的纹理保真与姿态一致性
- 适合虚拟试衣、数字人、手语生成等需高保真人像生成的应用
姿势引导人像生成(PGPIS)旨在根据指定姿势生成保留源图像身份与外观的人像,广泛应用于虚拟试衣、数字人、动画和手语生成。尽管基于扩散模型的最新方法已取得高质量结果,但通常依赖去噪过程中的隐式特征聚合,导致细粒度纹理保真度有限,且在不同姿势与源图像变化下难以保证生成一致性。为此,本文提出基于扩散模型的显式融合嵌入框架(FPDM),首次通过对比学习显式对齐源图-姿态融合嵌入与目标图像嵌入,并将学习到的融合嵌入作为生成条件。FPDM采用源增强姿态融合(SEPF)方法,结合图像-姿态融合(IPF)模块,学习与目标图像对齐的融合嵌入,并利用源外观、目标姿态及学习到的融合嵌入共同引导条件扩散模型。在DeepFashion基准和RWTH-PHOENIX-Weather 2014T数据集上的实验表明,该方法在定量与定性评估中均达到领先水平,消融实验证明显式融合嵌入对提升纹理保真度和跨姿态/源外观变化的一致性具有显著作用。
原文摘要 · Abstract (English)
Pose-Guided Person Image Synthesis (PGPIS) aims to generate human images in specified poses while preserving the identity and appearance of a source image. This technology facilitates diverse applications, including virtual try-on, digital avatars, animation, and sign language generation. Despite the high-quality results of recent diffusion-based PGPIS, these models typically depend on implicit feature aggregation within the denoising process. As a result, fine-grained texture preservation is limited, and even for the same identity, it is difficult to ensure consistent generation under variations in pose and source appearance. To address these limitations, we propose Fusion Embedding for PGPIS using a Diffusion Model (FPDM), the first framework that explicitly aligns fused source-pose embeddings with target image embeddings via contrastive learning, and subsequently employs the learned fusion embedding as a conditioning signal for generation. FPDM integrates an Image-Pose Fusion (IPF) module into our proposed Source-Enhanced Pose Fusion approach to learn a fusion embedding aligned with the target image. We then employ a conditional diffusion model guided by source appearance, target pose, and the learned fusion embedding. Experiments on the DeepFashion benchmark and the RWTH-PHOENIX-Weather 2014T dataset demonstrate competitive performance compared to existing methods in both quantitative and qualitative evaluations, with ablation studies confirming that explicit fusion embedding alignment substantially improves texture fidelity and consistency across pose and source appearance variations.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。