揭示了漂移模型与得分生成模型的本质联系,发现其运输方向几乎等同于得分匹配。
A Unified View of Score-Based and Drifting Models
- 用核方法非参数估计得分,实现数据与模型分布的匹配
- 高维下拉普拉斯核的残差可忽略,运输方向趋近于得分方向
- 为非参数生成模型提供理论支撑,适合研究生成机制的学者
漂移模型通过优化数据与模型分布间的核诱导均值偏移差异来训练单步生成器,实践中默认使用拉普拉斯核。在每个点上,该差异比较向邻近数据样本和邻近模型样本的加权位移,从而定义生成样本的传输方向。本文表明,漂移模型与基于得分的生成建模存在更紧密联系:对于高斯核,总体均值偏移场恰好等于高斯平滑后数据与模型分布得分(即梯度对数密度)的差值。这一恒等式源于Tweedie公式,将高斯平滑密度的得分与其条件均值关联,表明高斯核漂移本质上是针对平滑分布的得分匹配目标。更一般地,我们推导出径向核的精确分解:均值偏移等于得分场加上残差项。对于实际使用的拉普拉斯核,我们从理论上和实验上证明该残差在高维下可忽略,意味着实际使用的传输场几乎为得分驱动。结果揭示了与扩散模型的结构关联:两者均使用得分不匹配的传输方向,但漂移模型通过核估计非参数化实现,而扩散模型则用神经网络参数化学习。
原文摘要 · Abstract (English)
Drifting models train one-step generators by optimizing a kernel-induced mean-shift discrepancy between the data and model distributions, with Laplace kernels used by default in practice. At each point, this discrepancy compares the kernel-weighted displacement toward nearby data samples with the corresponding displacement toward nearby model samples, thereby defining a transport direction for generated samples. In this paper, we show that drifting is more closely connected to score-based generative modeling than it may first appear, establishing a precise link to the score-matching principle underlying diffusion models. For Gaussian kernels, the population mean-shift field exactly equals the difference between the scores (i.e., the gradient-log-densities) of the Gaussian-smoothed data and model distributions. This identity follows from Tweedie's formula, which links the score of a Gaussian-smoothed density to its conditional mean, and implies that Gaussian-kernel drifting is exactly a score-matching objective on smoothed distributions. More generally, we derive an exact decomposition for radial kernels in which mean shift equals a score-based field plus a residual term. For the practical Laplace kernel, we further show theoretically and empirically that this residual is negligible in high dimension, implying that the transport field used in practice is nearly score-based. Our results reveal a structural connection to diffusion models: both methods use score-mismatch transport directions, but drifting realizes the score nonparametrically through kernel-based estimates, whereas diffusion models learn it parametrically with neural networks.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。