用人体解析引导注意力,让换姿势时衣服细节不丢失。
TruePose: Human-Parsing-guided Attention Diffusion for Full-ID Preserving Pose Transfer
- 引入人体解析图指导扩散模型注意力,精准保留衣物纹理。
- 在换姿差异大时仍能保持面部与服装特征,优于13种基线方法。
- 适合时尚行业换装生成,对版权保护场景很实用。
姿态引导的人像合成(PGPIS)旨在保持源图像中人物身份的同时,采用目标姿态生成新图像(如骨骼结构)。尽管基于扩散模型的PGPIS方法在姿态变换中能较好保留面部特征,但往往难以准确维持源图像中的衣物细节,尤其在源姿态与目标姿态差异显著时更为严重,这在时尚产业中影响重大,因衣物风格的准确保留关乎版权保护。分析表明,该问题主要源于条件扩散模型的注意力模块未能充分捕捉和保留衣物图案。为此,本文提出人类解析引导的注意力扩散方法(TruePose),通过一个包含双相同UNet(TargetNet用于去噪,SourceNet用于提取源图嵌入)、人体解析引导融合注意力(HPFA)和CLIP引导注意力对齐(CAA)的孪生网络,自适应地将面部与衣物图案嵌入目标图像生成过程。在in-shop衣物检索基准和最新野外人体编辑数据集上的大量实验表明,该方法在保持面部和衣物外观方面显著优于13种基线模型。
原文摘要 · Abstract (English)
Pose-Guided Person Image Synthesis (PGPIS) generates images that maintain a subject's identity from a source image while adopting a specified target pose (e.g., skeleton). While diffusion-based PGPIS methods effectively preserve facial features during pose transformation, they often struggle to accurately maintain clothing details from the source image throughout the diffusion process. This limitation becomes particularly problematic when there is a substantial difference between the source and target poses, significantly impacting PGPIS applications in the fashion industry where clothing style preservation is crucial for copyright protection. Our analysis reveals that this limitation primarily stems from the conditional diffusion model's attention modules failing to adequately capture and preserve clothing patterns. To address this limitation, we propose human-parsing-guided attention diffusion, a novel approach that effectively preserves both facial and clothing appearance while generating high-quality results. We propose a human-parsing-aware Siamese network that consists of three key components: dual identical UNets (TargetNet for diffusion denoising and SourceNet for source image embedding extraction), a human-parsing-guided fusion attention (HPFA), and a CLIP-guided attention alignment (CAA). The HPFA and CAA modules can embed the face and clothes patterns into the target image generation adaptively and effectively. Extensive experiments on both the in-shop clothes retrieval benchmark and the latest in-the-wild human editing dataset demonstrate our method's significant advantages over 13 baseline approaches for preserving both facial and clothes appearance in the source image.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。