提出评估生成图像中身份保持能力的基准,验证持久身份表征的有效性。
Persistent Identity Preservation in Generative Image Models: A Benchmark and Evaluation System

- 对比输入上下文、可训练参数与持久身份层三种身份表示方法
- 发现迭代编辑等场景下身份漂移严重,持久身份可显著减少退化
- 适合关注人物/物体一致性的图像生成与编辑研究者
生成式图像模型虽能生成高质量图像并精准遵循指令,但在持续生成或编辑特定主体时仍面临身份漂移问题。现有方法将身份信息分别存储于输入上下文(GPT-Image-2、NB2)、可训练的主体特定参数(LoRA)或可复用的持久身份层(PHOTA IDENTITY)。本文系统评测这三类范式在主体驱动生成、编辑、修复及多主体场景下的表现,任务设计逐步加剧对身份保持的挑战。结果表明:当前生成基础模型的身份保持仍是独立短板——强图像质量与指令遵循不等于高身份保真度;身份退化在迭代编辑、小主体尺度、严重图像退化及多主体组合下更显著。引入持久身份层可有效缓解该问题,在不同基础模型上均提升身份保持能力,同时维持相近的指令遵循与感知图像质量。这说明身份并非仅由模型能力自然涌现,而是可作为独立于生成模型的持久主体知识进行构建。
原文摘要 · Abstract (English)
Generative image models can now produce high-quality images, follow complex instructions, and support precise edits, but they still struggle to preserve who or what is being depicted. When generating or editing images of a specific subject, identity may drift as the pose, expression, appearance, viewpoint, or surrounding scene changes. Existing subject-driven methods make fundamentally different choices about where identity is represented: through the input context (GPT-Image-2, NB2), as trainable subject-specific model parameters (LoRA), or as a persistent identity layer (PHOTA IDENTITY) reusable across generations and edits. We systematically benchmark these paradigms across subject-driven generation, editing, restoration, and multi-subject settings, with tasks designed to increasingly stress identity preservation. Our results show that identity preservation remains a distinct limitation of current generative foundation models: strong image quality and instruction following do not necessarily imply strong identity fidelity, and identity degradation becomes more pronounced under iterative edits, small subject scales, severe image degradation, and multi-subject composition. Persistent identity substantially reduces this degradation across generation, editing, and restoration, consistently improving identity preservation when applied to different foundation models while maintaining comparable instruction adherence and perceptual image quality. These results suggest that identity does not simply emerge from increasingly capable generative models, but can instead be represented as persistent subject knowledge that is composed independently with the underlying generative model.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。