arXiv:2510.14553cs.CV2025-10中稿 · ICLR被引 1

解决文本生成图像中人物身份漂移问题,无需提前知道所有场景。

Consistent text-to-image generation via scene de-contextualization

  • 通过反向消除图像生成模型中固有的场景与身份关联来稳定人物形象。
  • 在多个场景下生成图像时,身份保持率提升超过30%,且不损失场景多样性。
  • 无需训练、可逐场景使用,适合真实应用中场景未知或动态变化的场景。

一致的文本到图像生成旨在跨不同场景保持同一主体的身份一致性,但常因身份漂移现象而失败。现有方法通常假设需预先知晓所有目标场景,这一前提不切实际。本文揭示身份漂移的关键来源是主体与场景上下文之间的自然关联,称为场景上下文化,这是由于生成模型拟合海量自然图像分布所导致。我们形式化证明了该场景-身份关联的近普遍性,并推导出其强度的理论边界。基于此,提出一种新颖、高效、无需训练的提示嵌入编辑方法——场景去上下文化(SDeC),通过量化奇异值分解方向稳定性,自适应地重加权以抑制提示嵌入中的隐式场景-身份关联。关键优势在于支持单场景提示使用,无需提前获取所有目标场景。实验表明,SDeC显著提升身份保持能力,同时维持场景多样性。

原文摘要 · Abstract (English)

Consistent text-to-image (T2I) generation seeks to produce identity-preserving images of the same subject across diverse scenes, yet it often fails due to a phenomenon called identity (ID) shift. Previous methods have tackled this issue, but typically rely on the unrealistic assumption of knowing all target scenes in advance. This paper reveals that a key source of ID shift is the native correlation between subject and scene context, called scene contextualization, which arises naturally as T2I models fit the training distribution of vast natural images. We formally prove the near-universality of this scene-ID correlation and derive theoretical bounds on its strength. On this basis, we propose a novel, efficient, training-free prompt embedding editing approach, called Scene De-Contextualization (SDeC), that imposes an inversion process of T2I's built-in scene contextualization. Specifically, it identifies and suppresses the latent scene-ID correlation within the ID prompt's embedding by quantifying the SVD directional stability to adaptively re-weight the corresponding eigenvalues. Critically, SDeC allows for per-scene use (one scene per prompt) without requiring prior access to all target scenes. This makes it a highly flexible and general solution well-suited to real-world applications where such prior knowledge is often unavailable or varies over time. Experiments demonstrate that SDeC significantly enhances identity preservation while maintaining scene diversity.

文本生成图像身份一致去上下文化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。