arXiv:2409.06364cs.LG2024-09

条件扩散模型的似然值对输入文本不敏感,揭示了其本质特性仍不明。

What happens to diffusion model likelihood when your model is conditional?

  • 通过微分方程推导采样过程,实现精确似然计算。
  • 语音合成中似然值完全忽略文本输入,图像生成中也难区分相似提示。
  • 研究揭示条件扩散模型似然的不可靠性,适合关注模型可解释性的研究者。

扩散模型通过迭代去噪生成高质量数据,其采样过程基于随机微分方程(SDE),支持推理时的速度-质量权衡。利用微分方程采样还允许精确计算似然值,已被用于无条件模型排序和域外分类。然而,条件任务(如文生图、文生音)中的似然性质尚不明确。令人意外的是,本文发现文生音扩散模型的似然值对文本输入完全不敏感;文生图模型虽更具表达性,但仍无法区分混淆提示。结果表明,将扩散模型应用于条件任务暴露了其似然性质的不一致性,进一步证明了当前对扩散模型似然特性的理解仍处于未知状态。尽管条件扩散模型在最大化似然,但该似然值对条件输入的敏感度远低于预期。本研究为理解扩散模型似然提供了新视角。

原文摘要 · Abstract (English)

Diffusion Models (DMs) iteratively denoise random samples to produce high-quality data. The iterative sampling process is derived from Stochastic Differential Equations (SDEs), allowing a speed-quality trade-off chosen at inference. Another advantage of sampling with differential equations is exact likelihood computation. These likelihoods have been used to rank unconditional DMs and for out-of-domain classification. Despite the many existing and possible uses of DM likelihoods, the distinct properties captured are unknown, especially in conditional contexts such as Text-To-Image (TTI) or Text-To-Speech synthesis (TTS). Surprisingly, we find that TTS DM likelihoods are agnostic to the text input. TTI likelihood is more expressive but cannot discern confounding prompts. Our results show that applying DMs to conditional tasks reveals inconsistencies and strengthens claims that the properties of DM likelihood are unknown. This impact sheds light on the previously unknown nature of DM likelihoods. Although conditional DMs maximise likelihood, the likelihood in question is not as sensitive to the conditioning input as one expects. This investigation provides a new point-of-view on diffusion likelihoods.

扩散模型似然分析条件生成可解释性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。