arXiv:2503.15686cs.CV2025-03CVPR被引 11

用多焦点条件增强扩散模型,让人物图像生成更真实、身份一致。

Multi-focal Conditioned Latent Diffusion for Person Image Synthesis

  • 通过分离敏感区域特征,实现姿态无关的细节条件控制
  • 在DeepFashion数据集上保持身份一致性与外观真实性
  • 适合需要高保真人物图像生成与编辑的场景

潜空间扩散模型(LDM)在高分辨率图像生成中表现出色,广泛应用于姿态引导的人物图像合成(PGPIS),取得了良好效果。然而,LDM的压缩过程常导致细节退化,尤其在面部特征和衣物纹理等敏感区域。本文提出多焦点条件潜空间扩散(MCLD)方法,通过在这些敏感区域使用解耦的、姿态无关的特征进行条件控制,以缓解该问题。我们设计了多焦点条件聚合模块,有效融合面部身份与纹理特异性信息,显著提升生成图像的外观真实性和身份一致性。在DeepFashion数据集上,该方法展现出稳定的身份与外观生成能力,并支持灵活的人物图像编辑。代码已开源:https://github.com/jqliu09/mcld。

原文摘要 · Abstract (English)

The Latent Diffusion Model (LDM) has demonstrated strong capabilities in high-resolution image generation and has been widely employed for Pose-Guided Person Image Synthesis (PGPIS), yielding promising results. However, the compression process of LDM often results in the deterioration of details, particularly in sensitive areas such as facial features and clothing textures. In this paper, we propose a Multi-focal Conditioned Latent Diffusion (MCLD) method to address these limitations by conditioning the model on disentangled, pose-invariant features from these sensitive regions. Our approach utilizes a multi-focal condition aggregation module, which effectively integrates facial identity and texture-specific information, enhancing the model's ability to produce appearance realistic and identity-consistent images. Our method demonstrates consistent identity and appearance generation on the DeepFashion dataset and enables flexible person image editing due to its generation consistency. The code is available at https://github.com/jqliu09/mcld.

人物生成扩散模型姿态引导

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。