统一控制人类与动物图像生成,支持多实例重叠与复杂遮挡
UniMC: Taming Diffusion Transformer for Unified Keypoint-Guided Multi-Class Image Generation
- 基于DiT框架,用紧凑令牌融合类别、框和关键点信息
- 在290万实例数据集上实现多类重叠物体生成,遮挡下仍保持高精度
- 适合需要精细控制非刚性物体生成的研究者与开发者
尽管关键点引导的文本到图像扩散模型已取得显著进展,现有主流方法在控制除人类外更通用的非刚性物体(如动物)时仍面临挑战,且难以仅通过关键点控制生成多个重叠的人类与动物。这些问题源于可控方法的固有局限与缺乏合适数据集。为此,我们设计了基于DiT的统一框架UniMC,将实例级与关键点级条件整合为紧凑令牌,包含类别、边界框与关键点坐标等属性,克服了此前依赖骨架图作为条件而难以区分实例与类别的缺陷。同时,我们提出大规模高质量数据集HAIG-2.9M,涵盖78.6万张图像中的290万实例,包含人类与动物的关键点、边界框及细粒度描述,并经严格人工质检确保标注准确。大量实验证明,HAIG-2.9M质量高,UniMC在复杂遮挡与多类别场景下表现优异。
原文摘要 · Abstract (English)
Although significant advancements have been achieved in the progress of keypoint-guided Text-to-Image diffusion models, existing mainstream keypoint-guided models encounter challenges in controlling the generation of more general non-rigid objects beyond humans (e.g., animals). Moreover, it is difficult to generate multiple overlapping humans and animals based on keypoint controls solely. These challenges arise from two main aspects: the inherent limitations of existing controllable methods and the lack of suitable datasets. First, we design a DiT-based framework, named UniMC, to explore unifying controllable multi-class image generation. UniMC integrates instance- and keypoint-level conditions into compact tokens, incorporating attributes such as class, bounding box, and keypoint coordinates. This approach overcomes the limitations of previous methods that struggled to distinguish instances and classes due to their reliance on skeleton images as conditions. Second, we propose HAIG-2.9M, a large-scale, high-quality, and diverse dataset designed for keypoint-guided human and animal image generation. HAIG-2.9M includes 786K images with 2.9M instances. This dataset features extensive annotations such as keypoints, bounding boxes, and fine-grained captions for both humans and animals, along with rigorous manual inspection to ensure annotation accuracy. Extensive experiments demonstrate the high quality of HAIG-2.9M and the effectiveness of UniMC, particularly in heavy occlusions and multi-class scenarios.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。