分两阶段生成多人图像,确保人脸位置准确且身份不混淆。
Ar2Can: An Architect and an Artist Leveraging a Canvas for Multi-Human Generation
- 先规划布局再生成图像,分离空间位置与身份渲染
- 在测试集上人脸数量准确率和身份保留率显著提升
- 仅用合成数据训练,适合无真实多人图像的场景
尽管个性化图像生成取得进展,现有模型在生成多人场景时仍常出现人脸融合或身份丢失。本文提出Ar2Can,一种两阶段框架,将空间布局规划与身份渲染解耦。架构师(Architect)预测结构化布局,指定每个人的位置;艺术家(Artist)则基于空间引导的人脸匹配奖励,生成逼真图像,该奖励结合匈牙利算法的空间对齐与身份相似性。此方法确保人脸位于正确位置并忠实保留参考身份。我们开发了两种架构师变体,无缝集成于基于扩散的艺术家模型中,并通过组合奖励(计数准确性、图像质量、身份匹配)使用群体相对策略优化(GRPO)进行训练。在MultiHuman-Testbench上评估,Ar2Can在计数准确性和身份保留方面均实现显著提升,同时保持高感知质量。值得注意的是,本方法主要使用合成数据训练,无需真实多人图像。
原文摘要 · Abstract (English)
Despite recent advances in personalized image generation, existing models consistently fail to produce reliable multi-human scenes, often merging or losing facial identity. We present Ar2Can, a novel two-stage framework that disentangles spatial planning from identity rendering for multi-human generation. The Architect predicts structured layouts, specifying where each person should appear. The Artist then synthesizes photorealistic images, guided by a spatially-grounded face matching reward that combines Hungarian spatial alignment with identity similarity. This approach ensures faces are rendered at correct locations and faithfully preserve reference identities. We develop two Architect variants, seamlessly integrated with our diffusion-based Artist model. This is optimized via Group Relative Policy Optimization (GRPO) using compositional rewards for count accuracy, image quality, and identity matching. Evaluated on the MultiHuman-Testbench, Ar2Can achieves substantial improvements in both count accuracy and identity preservation, while maintaining high perceptual quality. Notably, our method achieves these results using primarily synthetic data, without requiring real multi-human images. Project page: https://qualcomm-ai-research.github.io/ar2can/.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。