arXiv:2410.16431cs.AI2024-10

用文本生成的图像分布距离衡量语义相似性,更贴近人类判断。

Conjuring Semantic Similarity

  • 以文本引发的图像分布差异度量语义相似性,而非传统重述关系。
  • 通过反向扩散SDE的Jeffreys散度实现可计算的相似性度量。
  • 适用于评估文生图模型性能,且能解释模型学到的语义表示。

样本表达间的语义相似性衡量其潜在‘含义’之间的距离,这些含义通常由文本表达表示。本文提出一种新方法:文本表达间的语义相似性不基于它们可被重述为的其他表达,而是基于其唤起的图像。虽然人类无法直接做到这一点,但生成模型可轻松生成并比较由文本提示引发的图像或其分布。因此,我们简单地将两个文本表达间的语义相似性定义为它们所诱导图像分布之间的距离,即‘召唤’(conjure)。通过选择由每个文本表达诱导的反向时间扩散随机微分方程(SDE)之间的Jeffreys散度,该距离可通过蒙特卡洛采样直接计算。该方法不仅与人工标注的语义相似性得分一致,还为文生图生成模型的评估提供了新路径,并提升了其学习表征的可解释性。

原文摘要 · Abstract (English)

The semantic similarity between sample expressions measures the distance between their latent 'meaning'. These meanings are themselves typically represented by textual expressions. We propose a novel approach whereby the semantic similarity among textual expressions is based not on other expressions they can be rephrased as, but rather based on the imagery they evoke. While this is not possible with humans, generative models allow us to easily visualize and compare generated images, or their distribution, evoked by a textual prompt. Therefore, we characterize the semantic similarity between two textual expressions simply as the distance between image distributions they induce, or 'conjure.' We show that by choosing the Jeffreys divergence between the reverse-time diffusion stochastic differential equations (SDEs) induced by each textual expression, this can be directly computed via Monte-Carlo sampling. Our method contributes a novel perspective on semantic similarity that not only aligns with human-annotated scores, but also opens up new avenues for the evaluation of text-conditioned generative models while offering better interpretability of their learnt representations.

语义相似性扩散模型图像生成可解释性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。