单次生成多人图像,保持身份一致且支持多级引导。
SIGMA-GEN: Structure and Identity Guided Multi-subject Assembly for Image Generation
- 基于结构与空间约束,单模型实现多身份保真生成。
- 在27,000张图像上训练,支持从框到像素的精细引导。
- 适用于需要精准控制多人布局的图像生成场景。
我们提出SIGMA-GEN,一种统一的多身份保真人像生成框架。与以往方法不同,SIGMA-GEN是首个支持单次生成、在结构与空间约束下保持多个主体身份一致的方法。其核心优势在于可接受多层级用户引导——从粗粒度的2D/3D框到像素级分割和深度图,仅需一个模型即可完成。为支持该能力,我们构建了SIGMA-SET27K,一个新型合成数据集,涵盖27,000张图像中超过10万唯一主体的身份、结构与空间信息。大量实验证明,SIGMA-GEN在身份保留、图像生成质量与速度方面均达到当前最优水平。
原文摘要 · Abstract (English)
We present SIGMA-GEN, a unified framework for multi-identity preserving image generation. Unlike prior approaches, SIGMA-GEN is the first to enable single-pass multi-subject identity-preserved generation guided by both structural and spatial constraints. A key strength of our method is its ability to support user guidance at various levels of precision -- from coarse 2D or 3D boxes to pixel-level segmentations and depth -- with a single model. To enable this, we introduce SIGMA-SET27K, a novel synthetic dataset that provides identity, structure, and spatial information for over 100k unique subjects across 27k images. Through extensive evaluation we demonstrate that SIGMA-GEN achieves state-of-the-art performance in identity preservation, image generation quality, and speed. Code and visualizations at https://oindrilasaha.github.io/SIGMA-Gen/
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。