arXiv:2602.18533cs.CV2026-02

用特征描述和发音结构可精准导航文生图模型中的身份空间。

Morphological Addressing of Identity Basins in Text-to-Image Diffusion Models

  • 通过特征词生成合成图像,训练LoRA实现身份定位。
  • 无照片仅凭描述词即达0.371的视觉一致性,最优达1.0。
  • 发音规律能生成独特且一致的视觉形象,适合创意设计者。

我们证明了形态压力在文本到图像生成流程中创造了多层级的可导航梯度。研究1显示,在Stable Diffusion 1.5中,仅用铂金色长发、美人痣、1950年代魅力等形态描述词,无需目标姓名或照片即可导航身份基域。通过自蒸馏循环(用描述词生成合成图像并训练LoRA)实现了弧面相似性一致收敛。训练后的LoRA构建了局部坐标系,不仅塑造目标身份,还产生反向效果:极大远离条件导致基模型出现诡异结构崩溃,而搭载LoRA的模型则生成‘恐怖谷’式输出——看似合理但精确错误。研究2扩展至提示级形态学,基于音形理论生成200个新造无意义词(如cr-、sn-、-oid、-ax),发现含音形特征的候选词显著提升视觉一致性(均值Purity@1 = 0.371 vs. 0.209,p<0.00001,Cohen's d=0.55)。三个词——snudgeoid、crashax、broomix——在零训练数据污染下达到完美视觉一致性(Purity@1=1.0),各自生成独立且连贯的视觉身份。两组研究共同表明,无论特征描述还是提示级语音结构,都能在扩散模型潜空间中形成系统性导航梯度。我们记录了身份基域的相变、CFG无关的身份稳定性,以及子音素音律模式引发的新视觉概念。

原文摘要 · Abstract (English)

We demonstrate that morphological pressure creates navigable gradients at multiple levels of the text-to-image generative pipeline. In Study~1, identity basins in Stable Diffusion 1.5 can be navigated using morphological descriptors -- constituent features like platinum blonde,'' beauty mark,'' and 1950s glamour'' -- without the target's name or photographs. A self-distillation loop (generating synthetic images from descriptor prompts, then training a LoRA on those outputs) achieves consistent convergence toward a specific identity as measured by ArcFace similarity. The trained LoRA creates a local coordinate system shaping not only the target identity but also its inverse: maximal away-conditioning produces eldritch'' structural breakdown in base SD1.5, while the LoRA-equipped model produces ``uncanny valley'' outputs -- coherent but precisely wrong. In Study~2, we extend this to prompt-level morphology. Drawing on phonestheme theory, we generate 200 novel nonsense words from English sound-symbolic clusters (e.g., \emph{cr-}, \emph{sn-}, \emph{-oid}, \emph{-ax}) and find that phonestheme-bearing candidates produce significantly more visually coherent outputs than random controls (mean Purity@1 = 0.371 vs.\ 0.209, p<0.00001p < 0.00001 p<0.00001, Cohen's d=0.55d = 0.55 d=0.55). Three candidates -- \emph{snudgeoid}, \emph{crashax}, and \emph{broomix} -- achieve perfect visual consistency (Purity@1 = 1.0) with zero training data contamination, each generating a distinct, coherent visual identity from phonesthetic structure alone. Together, these studies establish that morphological structure -- whether in feature descriptors or prompt-level phonological form -- creates systematic navigational gradients through diffusion model latent spaces. We document phase transitions in identity basins, CFG-invariant identity stability, and novel visual concepts emerging from sub-lexical sound patterns.

文生图身份导航音形理论扩散模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。