用语义锚定让扩散模型稳定个性化生成用户特定图像
Semantic Anchoring for Robust Personalization in Text-to-Image Diffusion Models
- 通过语义锚定将新概念引导至预训练分布中
- 在少量参考图下同时提升主体保真度与图文对齐
- 适合需要快速个性化的图像生成应用
文本到图像扩散模型在生成多样且逼真的图像方面取得了显著进展,但在个性化方面仍存在挑战:仅用少量参考图像适应预训练模型以描绘用户特定主体。核心难题在于,在有限参考图像下学习新视觉概念的同时,需保留预训练的语义先验以维持图文对齐。当模型关注主体保真度时,易过拟合有限参考图,无法利用预训练分布;而强调先验保持则会阻碍学习个性化特征。为此,我们提出基于语义锚定的个性化方法,通过将新概念锚定在其对应常见概念的分布上,实现稳定适应。该方法将个性化建模为在频繁概念引导下学习罕见概念的过程,使模型在扩展预训练分布至个性化区域的同时,保持其语义结构。实验与消融研究证明,该方法在主体保真度和图文对齐上均优于基线,具有更强鲁棒性与有效性。
原文摘要 · Abstract (English)
Text-to-image diffusion models have achieved remarkable progress in generating diverse and realistic images from textual descriptions. However, they still struggle with personalization, which requires adapting a pretrained model to depict user-specific subjects from only a few reference images. The key challenge lies in learning a new visual concept from a limited number of reference images while preserving the pretrained semantic prior that maintains text-image alignment. When the model focuses on subject fidelity, it tends to overfit the limited reference images and fails to leverage the pretrained distribution. Conversely, emphasizing prior preservation maintains semantic consistency but prevents the model from learning new personalized attributes. Building on these observations, we propose the personalization process through a semantic anchoring that guides adaptation by grounding new concepts in their corresponding distributions. We therefore reformulate personalization as the process of learning a rare concept guided by its frequent counterpart through semantic anchoring. This anchoring encourages the model to adapt new concepts in a stable and controlled manner, expanding the pretrained distribution toward personalized regions while preserving its semantic structure. As a result, the proposed method achieves stable adaptation and consistent improvements in both subject fidelity and text-image alignment compared to baseline methods. Extensive experiments and ablation studies further demonstrate the robustness and effectiveness of the proposed anchoring strategy.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。