arXiv:2512.23245cs.CV2025-12

无需训练即可保持角色一致性,同时精准响应每张图的文本提示。

ASemConsist: Adaptive Semantic Feature Control for Training-Free Identity-Consistent Generation

  • 通过选择性修改文本嵌入,实现身份一致与提示对齐的平衡。
  • 在SD3.5和FLUX架构上均超越现有方法,提升图像生成质量。
  • 适合需要高保真角色生成的视频创作、虚拟人设计等场景。

近期的文本到图像扩散模型在视觉质量和文本对齐方面取得了显著进展,但在跨场景生成连贯角色图像时仍面临挑战。现有方法常在身份一致性和单图提示对齐之间存在权衡。本文提出AsemConsist框架,通过选择性文本嵌入修改,实现无需训练的身份一致性保持,且不损害单图提示对齐。我们分析了填充嵌入的语义结构,发现多编码器骨干网络中仅保留提示相关语义的填充嵌入可有效作为语义容器。基于此,我们选择性地将每张图的语义注入此类嵌入,并抑制无关语义成分。此外,提出自适应特征共享策略,自动评估身份特异性,仅对模糊提示施加约束。最后,提出统一评估指标SeeSaw,衡量身份一致性与单图对齐的平衡性,并评估生成图像中身份与提示的反映程度。实验表明,该方法在SD3.5和FLUX骨干网络上均优于现有竞争方法。

原文摘要 · Abstract (English)

Recent text-to-image diffusion models have significantly improved visual quality and text alignment. However, generating a sequence of images while preserving consistent character identity across diverse scenes remains challenging. Existing methods often face a trade-off between maintaining identity consistency and per-image prompt alignment. In this paper, we introduce AsemConsist, a framework that resolves this trade-off through selective text embedding modification, enabling consistent identity preservation without degrading per-image prompt alignment. We further analyze the semantic structure of padding embeddings and find that, in multi-encoder backbones, only padding embeddings that retain prompt-related semantics can effectively serve as semantic containers. Based on this observation, we selectively inject per-image semantics into such padding embeddings while suppressing prompt-irrelevant components. Additionally, we propose an adaptive feature-sharing strategy that automatically evaluates identity specificity and selectively applies constraints only to ambiguous identity prompts. Finally, we propose a unified evaluation metric called SeeSaw, which measures the balance between identity consistency and per-image alignment while evaluating whether identity and per-image prompts are equally reflected in generated images. Our method demonstrates superior performance over existing competitors when built upon SD3.5 and FLUX backbones, highlighting its effectiveness across different architectures and text encoders. Project page: https://minjung-s.github.io/asemconsist

图像生成身份一致扩散模型无训练

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。