arXiv:2608.20336cs.CV2026-08

让多人图像生成更准,能精准分配每个人的位置和身份。

WithEveryone: Unified Planning and Identity Grounding for Group Image Generation

论文配图:WithEveryone: Unified Planning and Identity Grounding for Group Image Generation
图 1 · 摘自论文原文
  • 用地址化标记注入身份,预测身份-布局计划并渲染为视觉条件。
  • 人脸相似度达0.499,复制粘贴伪影从0.169降至0.055。
  • 可覆盖97.3%请求身份,重复率仅2.8%,适合大规模群像生成。

当场景需包含多名指定人物时,保持身份一致性的图像生成变得不可靠。模型不仅需保留每个身份,还需将每处引用绑定到特定人物与位置,且训练阶段的身份损失需在多个噪声预测人脸间建立对应关系。我们提出WithEveryone,一个支持最多十名参考身份的统一框架。该框架将每个选定身份作为地址化标记注入,预测结构化的身份-布局计划,并将其渲染为视觉条件。其核心目标——布局接地身份损失,利用标注的人脸区域直接监督预期身份,避免不稳定的基于嵌入的脸部匹配;身份表示强制机制则在图像合成前训练每个身份的独立预测。在身份互斥基准上,WithEveryone实现了最高的目标-上下文身份相似度,将人脸相似度从GPT-Image-2的0.462提升至0.499,同时将复制粘贴伪影从0.169降低至0.055。此外,它能覆盖97.3%的请求身份,重复率仅为2.8%。结果表明,显式的身份-布局对齐使身份保持生成可扩展至更大群体,而无需依赖直接的参考人脸复制。

原文摘要 · Abstract (English)

Identity-preserving image generation becomes increasingly unreliable when a scene must contain many specified people. Beyond retaining each identity, the model must bind every reference to a distinct person and location, while training-time identity losses must establish correspondence among several noisy predicted faces. We introduce WithEveryone, a unified framework for generating group images up to ten reference identities. WithEveryone injects each selected identity as an addressed token, predicts a structured identity--layout plan, and renders the plan as a visual condition. Its key objective, Layout-Grounded ID Loss, uses annotated face regions to supervise the intended identities directly, avoiding unstable embedding-based face matching; ID Representation Forcing additionally trains a prediction for each identity before image synthesis. On an identity-disjoint benchmark, WithEveryone achieves the highest target-context identity similarity, improving face similarity from 0.462 for GPT-Image-2 to 0.499, while reducing copy-paste artifacts from 0.169 to 0.055. It further covers 97.3\% of the requested identities with a duplicate rate of only 2.8\%. These results show that explicit identity--layout grounding enables identity-preserving generation to scale to larger groups without relying on direct reference-face copying.

图像生成身份保持多人群像

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。