通过提示干预分析扩散模型中概念生成的时间动态,揭示何时概念锁定轨迹。
Temporal Concept Dynamics in Diffusion Models via Prompt-Conditioned Interventions
- 提出无需训练的PCI框架,通过概念插入成功率分析时间动态。
- 发现不同阶段对特定概念的形成更有利,同一概念类型内存在差异。
- 为文本驱动图像编辑提供有效时机建议,无需访问模型内部。
扩散模型通常通过最终输出评估,逐步将随机噪声还原为有意义的图像。然而生成过程沿轨迹展开,分析这一动态过程对理解模型可控性、可靠性和可预测性至关重要。本文探讨:噪声何时转化为特定概念(如年龄)并锁定去噪路径?提出PCI(提示条件干预)框架,用于分析扩散时间中的概念动态。核心思想是概念插入成功率(CIS),即在特定时间步插入概念后其在最终图像中保留的概率,用以刻画概念形成的时序特性。应用于多个前沿文生图扩散模型及广泛概念类别,结果揭示不同模型在概念形成上存在多样化的时序行为,即使同一概念类型也存在不同有利阶段。这些发现为文本驱动图像编辑提供了可操作洞见,指明干预的最佳时机,且无需模型内部信息或训练,实现比强基线更优的语义准确性和内容保真度。代码已公开:https://adagorgun.github.io/PCI-Project/
原文摘要 · Abstract (English)
Diffusion models are usually evaluated by their final outputs, gradually denoising random noise into meaningful images. Yet, generation unfolds along a trajectory, and analyzing this dynamic process is crucial for understanding how controllable, reliable, and predictable these models are in terms of their success/failure modes. In this work, we ask the question: when does noise turn into a specific concept (e.g., age) and lock in the denoising trajectory? We propose PCI (Prompt-Conditioned Intervention) to study this question. PCI is a training-free and model-agnostic framework for analyzing concept dynamics through diffusion time. The central idea is the analysis of Concept Insertion Success (CIS), defined as the probability that a concept inserted at a given timestep is preserved and reflected in the final image, offering a way to characterize the temporal dynamics of concept formation. Applied to several state-of-the-art text-to-image diffusion models and a broad taxonomy of concepts, PCI reveals diverse temporal behaviors across diffusion models, in which certain phases of the trajectory are more favorable to specific concepts even within the same concept type. These findings also provide actionable insights for text-driven image editing, highlighting when interventions are most effective without requiring access to model internals or training, and yielding quantitatively stronger edits that achieve a balance of semantic accuracy and content preservation than strong baselines. Code is available at: https://adagorgun.github.io/PCI-Project/
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。