用草图+文字/图片精准控制人体生成,灵活又准确。
ComposeAnyone: Controllable Layout-to-Human Generation with Decoupled Multimodal Conditions
- 草图用色块形状定义布局,支持文本或参考图分别控制各部位。
- 在多数据集上生成图像与布局、描述、参考图对齐度更高。
- 适合需要精细人体设计的时尚、影视等领域应用。
基于扩散模型,多模态图像生成取得显著进展。其中,人体图像生成成为有前景的技术,有望革新时尚设计流程。然而,现有方法多局限于纯文本或参考图驱动的人体生成,难以满足日益复杂的需求。为解决灵活性与精度不足的问题,本文提出ComposeAnyone,一种解耦多模态条件的可控布局到人体生成方法。该方法允许使用文本或参考图像分别控制手绘人体布局中的任意部分,并在生成过程中无缝融合。手绘布局采用椭圆、矩形等色块几何形状,易于绘制,提供更灵活、易用的空间布局方式。此外,我们构建了ComposeHuman数据集,为每张人体图像提供解耦的文本与参考图像标注,涵盖不同身体组件,支持更广泛的人体图像生成任务。在多个数据集上的大量实验表明,ComposeAnyone生成的人体图像在布局、文本描述和参考图像对齐方面表现更优,展现出强大的多任务能力与可控性。
原文摘要 · Abstract (English)
Building on the success of diffusion models, significant advancements have been made in multimodal image generation tasks. Among these, human image generation has emerged as a promising technique, offering the potential to revolutionize the fashion design process. However, existing methods often focus solely on text-to-image or image reference-based human generation, which fails to satisfy the increasingly sophisticated demands. To address the limitations of flexibility and precision in human generation, we introduce ComposeAnyone, a controllable layout-to-human generation method with decoupled multimodal conditions. Specifically, our method allows decoupled control of any part in hand-drawn human layouts using text or reference images, seamlessly integrating them during the generation process. The hand-drawn layout, which utilizes color-blocked geometric shapes such as ellipses and rectangles, can be easily drawn, offering a more flexible and accessible way to define spatial layouts. Additionally, we introduce the ComposeHuman dataset, which provides decoupled text and reference image annotations for different components of each human image, enabling broader applications in human image generation tasks. Extensive experiments on multiple datasets demonstrate that ComposeAnyone generates human images with better alignment to given layouts, text descriptions, and reference images, showcasing its multi-task capability and controllability.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。