arXiv:2411.12872cs.CVcs.AI2024-11被引 2

让文字生成更精准的人体姿态图像,提升可控性与质量。

From Text to Pose to Image: Improving Diffusion Model Control and Quality

  • 用文本生成姿态的模型解决描述多样性问题
  • 新采样算法与适配器使姿态还原度更高
  • 适合需要精确人体控制的创作场景

近两年来,文本到图像的扩散模型广受欢迎。随着其质量和应用增长,输出控制成为关键挑战。除提示工程外,通过图像风格、深度图或关键点等模态进行条件控制是有效方法,如ControlNets或适配器。在将这些方法用于控制文本到图像扩散模型中的人体姿态时,面临两大挑战:一是生成符合广泛语义描述的姿态,此前方法依赖(标题, 姿态)数据集搜索;二是指定姿态条件下生成高审美且高姿态保真度的图像。本文通过引入文本到姿态(T2P)生成模型和新采样算法,以及融合更多姿态关键点的新姿态适配器,首次实现文本→姿态→图像的生成框架,显著提升扩散模型中的姿态控制能力。所有模型及实验代码已公开于https://github.com/clement-bonnet/text-to-pose。

原文摘要 · Abstract (English)

In the last two years, text-to-image diffusion models have become extremely popular. As their quality and usage increase, a major concern has been the need for better output control. In addition to prompt engineering, one effective method to improve the controllability of diffusion models has been to condition them on additional modalities such as image style, depth map, or keypoints. This forms the basis of ControlNets or Adapters. When attempting to apply these methods to control human poses in outputs of text-to-image diffusion models, two main challenges have arisen. The first challenge is generating poses following a wide range of semantic text descriptions, for which previous methods involved searching for a pose within a dataset of (caption, pose) pairs. The second challenge is conditioning image generation on a specified pose while keeping both high aesthetic and high pose fidelity. In this article, we fix these two main issues by introducing a text-to-pose (T2P) generative model alongside a new sampling algorithm, and a new pose adapter that incorporates more pose keypoints for higher pose fidelity. Together, these two new state-of-the-art models enable, for the first time, a generative text-to-pose-to-image framework for higher pose control in diffusion models. We release all models and the code used for the experiments at https://github.com/clement-bonnet/text-to-pose.

扩散模型姿态生成文本生成图像控制

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。