用知识库与动态掩码提升人体生成的姿势准确性和图像质量
KB-DMGen: Knowledge-Based Global Guidance and Dynamic Pose Masking for Human Image Generation
- 引入视觉代码本提供全局姿态指导,结合动态掩码实现局部精细控制
- 在HumanArt数据集上达到新SOTA,AP和CAP指标显著提升
- 适合关注人体生成中姿态与图像质量平衡的研究者
近期基于扩散模型的人体图像生成方法在姿态先验等控制信号下取得显著进展。在人体图像生成中,准确的姿态与连贯的视觉质量同样重要。然而,多数现有方法仅关注姿态准确性,常以牺牲图像质量为代价。为此,我们提出知识基全局引导与动态姿态掩码的人体图像生成框架(KB-DMGen)。知识库(KB)作为视觉代码本,基于输入文本相关的视觉特征提供粗粒度全局引导,在保持图像质量的同时提升姿态准确性;动态姿态掩码(DM)则提供细粒度局部控制,进一步增强姿态精确性。通过在扩散过程不同阶段注入KB与DM,该框架在不损害图像质量的前提下,通过全局与局部双重控制提升姿态准确性。实验表明,KB-DMGen在HumanArt数据集上实现了新的最先进水平,其AP与CAP指标均显著优于现有方法。项目页面与代码已公开于https://lushbng.github.io/KBDMGen。
原文摘要 · Abstract (English)
Recent methods using diffusion models have made significant progress in Human Image Generation (HIG) with various control signals such as pose priors. In HIG, both accurate human poses and coherent visual quality are crucial for image generation. However, most existing methods mainly focus on pose accuracy while neglecting overall image quality, often improving pose alignment at the cost of image quality. To address this, we propose Knowledge-Based Global Guidance and Dynamic pose Masking for human image Generation (KB-DMGen). The Knowledge Base (KB), implemented as a visual codebook, provides coarse, global guidance based on input text-related visual features, improving pose accuracy while maintaining image quality, while the Dynamic pose Mask (DM) offers fine-grained local control to enhance precise pose accuracy. By injecting KB and DM at different stages of the diffusion process, our framework enhances pose accuracy through both global and local control without compromising image quality. Experiments demonstrate the effectiveness of KB-DMGen, achieving new state-of-the-art results in terms of AP and CAP on the HumanArt dataset. The project page and code are available at https://lushbng.github.io/KBDMGen.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。