arXiv:2410.00483cs.CVcs.AI2024-10被引 1

用掩码控制生成人物姿态,单图就能调姿势。

MCGM: Mask Conditional Text-to-Image Generative Model

  • 通过掩码嵌入注入,让扩散模型按指定姿态生成图像。
  • 在单图基础上实现多主体姿态可控生成,效果优于原Break-a-scene模型。
  • 适合需要精确控制人物姿态的创意设计、动画制作场景。

生成模型的最新进展已彻底改变人工智能领域,实现了高度逼真且细节丰富的图像生成。本文提出一种新型的掩码条件文本到图像生成模型(MCGM),利用条件扩散模型生成特定姿态的图像。该模型基于Break-a-scene [1] 模型在单张多主体图像上生成新场景的成功经验,并引入掩码嵌入注入机制,实现对生成过程的额外控制。通过这一机制,MCGM 能够从单张图像中学习一个或多个主体的特定姿态,并根据用户需求灵活生成符合预设掩码条件的图像。经大量实验与评估,验证了该模型在生成高质量、满足掩码约束图像方面的有效性,同时提升了现有Break-a-scene生成模型的表现。

原文摘要 · Abstract (English)

Recent advancements in generative models have revolutionized the field of artificial intelligence, enabling the creation of highly-realistic and detailed images. In this study, we propose a novel Mask Conditional Text-to-Image Generative Model (MCGM) that leverages the power of conditional diffusion models to generate pictures with specific poses. Our model builds upon the success of the Break-a-scene [1] model in generating new scenes using a single image with multiple subjects and incorporates a mask embedding injection that allows the conditioning of the generation process. By introducing this additional level of control, MCGM offers a flexible and intuitive approach for generating specific poses for one or more subjects learned from a single image, empowering users to influence the output based on their requirements. Through extensive experimentation and evaluation, we demonstrate the effectiveness of our proposed model in generating high-quality images that meet predefined mask conditions and improving the current Break-a-scene generative model.

文本生成图像姿态控制扩散模型掩码条件

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。