用3D人体模型控制生成图像中人物的形状与姿态,提升多样性与稳定性。
Controlling Human Shape and Pose in Text-to-Image Diffusion Models via Domain Adaptation
- 通过合成数据训练控制网络,结合领域自适应保持图像质量。
- 相比2D姿态控制,生成的人体形状与姿态更丰富多样。
- 适合需要精准人体控制的动画生成、虚拟试衣等应用。
我们提出一种在预训练文本到图像扩散模型中实现人体形状与姿态条件控制的方法,采用3D人体参数化模型SMPL。微调这些扩散模型以适应新条件通常需大量高质量标注数据,而合成数据生成可更高效地解决这一问题。然而,合成数据存在领域差距和场景多样性不足的问题,可能损害预训练模型的视觉保真度。为此,我们设计了一种领域自适应技术:将合成训练的条件信息隔离于无分类器引导向量中,并与另一控制网络组合,使生成图像适配输入域。为实现对SMPL的控制,我们在合成渲染的人体数据集SURREAL上微调基于ControlNet的架构,并在生成时应用该领域自适应方法。实验表明,本模型在保持高视觉保真度的同时,比基于2D姿态的ControlNet生成更多样化的形状与姿态,且提升稳定性,适用于人体动画等下游任务。
原文摘要 · Abstract (English)
We present a methodology for conditional control of human shape and pose in pretrained text-to-image diffusion models using a 3D human parametric model (SMPL). Fine-tuning these diffusion models to adhere to new conditions requires large datasets and high-quality annotations, which can be more cost-effectively acquired through synthetic data generation rather than real-world data. However, the domain gap and low scene diversity of synthetic data can compromise the pretrained model's visual fidelity. We propose a domain-adaptation technique that maintains image quality by isolating synthetically trained conditional information in the classifier-free guidance vector and composing it with another control network to adapt the generated images to the input domain. To achieve SMPL control, we fine-tune a ControlNet-based architecture on the synthetic SURREAL dataset of rendered humans and apply our domain adaptation at generation time. Experiments demonstrate that our model achieves greater shape and pose diversity than the 2d pose-based ControlNet, while maintaining the visual fidelity and improving stability, proving its usefulness for downstream tasks such as human animation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。