用扩散模型生成更稳定多样的文字驱动人体姿态骨架
GUNet: A Graph Convolutional Network United Diffusion Model for Stable and Diversity Pose Generation
- 融合图卷积网络的去噪模型,学习人体骨架空间关系
- 解耦关节点并引入交叉注意力,提升生成多样性
- 适合需要可控姿态生成的图像合成任务
姿态骨架图在可控图像生成中具有重要参考价值。为丰富骨架来源,近期工作尝试基于自然语言生成姿态骨架,但现有方法多依赖GAN,难以同时保证生成姿态的结构正确性、多样性与美观性。为此,我们提出以GUNet为核心模型的PoseDiffusion框架,首个基于扩散模型的生成框架,并包含多个基于Stable Diffusion微调的变体。该框架具备多项优势:1)结构正确性:GUNet作为去噪模型,引入骨架信息训练,能有效学习人体骨架的空间关系;2)生成多样性:将骨架关键点解耦并分别建模,通过交叉注意力引入文本条件。实验表明,PoseDiffusion在文本驱动的姿态骨架生成上,稳定性与多样性均优于现有最先进方法。定性分析进一步验证其在Stable Diffusion中实现可控生成的优越性。
原文摘要 · Abstract (English)
Pose skeleton images are an important reference in pose-controllable image generation. In order to enrich the source of skeleton images, recent works have investigated the generation of pose skeletons based on natural language. These methods are based on GANs. However, it remains challenging to perform diverse, structurally correct and aesthetically pleasing human pose skeleton generation with various textual inputs. To address this problem, we propose a framework with GUNet as the main model, PoseDiffusion. It is the first generative framework based on a diffusion model and also contains a series of variants fine-tuned based on a stable diffusion model. PoseDiffusion demonstrates several desired properties that outperform existing methods. 1) Correct Skeletons. GUNet, a denoising model of PoseDiffusion, is designed to incorporate graphical convolutional neural networks. It is able to learn the spatial relationships of the human skeleton by introducing skeletal information during the training process. 2) Diversity. We decouple the key points of the skeleton and characterise them separately, and use cross-attention to introduce textual conditions. Experimental results show that PoseDiffusion outperforms existing SoTA algorithms in terms of stability and diversity of text-driven pose skeleton generation. Qualitative analyses further demonstrate its superiority for controllable generation in Stable Diffusion.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。