将分数匹配与最大似然、EM算法统一,解决混合线性回归参数估计问题。
Connecting Score Matching, Maximum Likelihood, and Expectation-Maximization in Mixed Linear Regression

- 通过扩散过程连接分数匹配与似然函数,分离统计保证与优化信号。
- 在低噪声下,误差收敛至最大似然估计的高斯极限,证明一致性。
- 揭示分数匹配梯度与EM算法的内在关联,适合统计学习研究者参考。
我们研究了在未知混合权重下响应变量的方差保持扩散过程在混合线性回归(MLR)中的应用。分析表明,分数匹配的统计保证可与损失几何和固定扩散噪声水平下的优化信号相分离。KL散度将沿扩散路径积分的去噪分数匹配目标与似然及终端偏差联系起来。在温和正则条件下,终端调度下估计量收敛至真实参数,其缩放误差收敛至最大似然估计的高斯极限。在固定扩散噪声水平下,我们推导出分数匹配损失与交叉熵及期望最大化(EM)算子之间的分解关系。该分解导出了一个与EM相关的低噪声梯度展开式,包含潜在方差的修正项。在高噪声极限下,进一步刻画了各向同性协方差下该极限损失上的梯度下降行为。沿固定信噪比射线,分数匹配失衡梯度与潜在方差项逐点趋于零。数值实验验证了理论发现与统计保证。
原文摘要 · Abstract (English)
We study variance-preserving diffusion of the response in mixed linear regression (MLR) with unknown mixing weights. Our analysis separates the statistical guarantees of score matching from the loss geometry and optimization signal at a fixed diffusion noise level. The KL divergence links the denoising score matching objective integrated over the diffusion path with the likelihood and a terminal discrepancy. Under mild regularity conditions and terminal schedule, the resulting estimator converges up to the ground truth parameters of MLR, and its scaled error converges to the Gaussian limit of the maximum-likelihood estimator. At a fixed scale of the diffusion noise level, we derive a decomposition linking the score matching loss to cross-entropy and Expectation-Maximization (EM) operators. This decomposition yields an EM-related low-noise gradient expansion with additional correction terms of latent variance. In the high-noise limit, we further characterize gradient descent on this limiting loss under isotropic covariance. Along fixed high signal-to-noise ratio rays, the score matching imbalance gradient and the latent-variance term tend to zero pointwise. Numerical experiments illustrate our theoretical findings and statistical guarantees.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。