arXiv:2603.26078cs.CVcs.AI2026-03中稿 · CVPR被引 1

测试多人物生成时身份混淆问题,发现模型复杂度上升后身份丢失严重。

When Identities Collapse: A Stress-Test Benchmark for Multi-Subject Personalization

  • 构建75个提示的压测基准,覆盖不同人数和交互难度
  • 发现10人场景下身份丢失率接近100%,传统评估方法失效
  • 提出新指标SCR,能准确捕捉局部身份混淆问题

基于主体的文本到图像扩散模型在保持单一身份方面表现优异,但在生成多个互动主体时仍面临严峻挑战。现有评估多依赖全局CLIP指标,对局部身份混淆不敏感,无法反映多主体纠缠的严重性。本文揭示当前模型存在普遍的“可扩展性幻觉”:在简单布局中生成2-4个主体表现良好,但当增至6-10个主体或涉及复杂物理交互时,身份会灾难性崩溃。为此,我们构建了包含75个提示的压测基准,涵盖不同主体数量与交互难度(中性、遮挡、交互)。实验表明,标准CLIP指标常给身份已崩溃但语义正确的图像高分(如生成通用克隆)。为此,我们提出基于DINOv2结构先验的主体坍塌率(SCR)指标,严格惩罚局部注意力泄漏与同质化。对MOSAIC、XVerse、PSR等前沿模型的评估显示,随着场景复杂度提升,身份保真度急剧下降,10主体时SCR接近100%。我们溯源发现该崩溃源于全局注意力路由中的语义捷径,凸显未来生成架构亟需显式物理解耦。

原文摘要 · Abstract (English)

Subject-driven text-to-image diffusion models have achieved remarkable success in preserving single identities, yet their ability to compose multiple interacting subjects remains largely unexplored and highly challenging. Existing evaluation protocols typically rely on global CLIP metrics, which are insensitive to local identity collapse and fail to capture the severity of multi-subject entanglement. In this paper, we identify a pervasive "Illusion of Scalability" in current models: while they excel at synthesizing 2-4 subjects in simple layouts, they suffer from catastrophic identity collapse when scaled to 6-10 subjects or tasked with complex physical interactions. To systematically expose this failure mode, we construct a rigorous stress-test benchmark comprising 75 prompts distributed across varying subject counts and interaction difficulties (Neutral, Occlusion, Interaction). Furthermore, we demonstrate that standard CLIP-based metrics are fundamentally flawed for this task, as they often assign high scores to semantically correct but identity-collapsed images (e.g., generating generic clones). To address this, we introduce the Subject Collapse Rate (SCR), a novel evaluation metric grounded in DINOv2's structural priors, which strictly penalizes local attention leakage and homogenization. Our extensive evaluation of state-of-the-art models (MOSAIC, XVerse, PSR) reveals a precipitous drop in identity fidelity as scene complexity grows, with SCR approaching 100% at 10 subjects. We trace this collapse to the semantic shortcuts inherent in global attention routing, underscoring the urgent need for explicit physical disentanglement in future generative architectures.

身份保持多主体生成评估基准扩散模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。