arXiv:2607.01895cs.LGmath.OC2026-07

比较了高维下两种密度比估计方法的性能,发现样本多时变分法更优,样本少时谱法更稳定。

Regularized Variational and Spectral Log-Density-Ratio Estimation in the Gaussian Location Model

论文配图:Regularized Variational and Spectral Log-Density-Ratio Estimation in the Gaussian Location Model
图 1 · 摘自论文原文
  • 用正则化变分法和谱法估计高斯位置模型中的密度比
  • 样本充足时变分法风险更低,样本少时谱法方差更小
  • 适用于高维统计推断与特征学习场景

我们研究在具有共同协方差矩阵的高斯位置模型中,使用岭正则化的对数密度比估计。通过仿射不变性,模型可写为 q ~ N(0, I),p ~ N(Δ, I),采用线性特征表示,其中 Δ 为均值向量。变分估计器是带平方 L2 惩罚的经验 KL 对数归一化拟合,而最近提出的谱估计器将单一变分问题替换为一系列岭正则化的最小二乘问题。当样本数和维度同时趋于无穷且比值固定时,我们推导出高维确定性渐近等价形式。正则化变分极限由凸高斯极小极大定理(CGMT)导出的标量熵最小化问题刻画,而正则化谱极限则来自两个独立高斯样本协方差矩阵加权和的再生核等价公式。利用这些公式比较总体风险,实验聚焦于固定信号-维度比扫描和最优正则化选择。结论表明:在观测数较多时,设定正确的变分估计器风险更小;而在观测较少时,基于协方差构造的谱估计器因方差更低而更优。此外,我们还研究了核范数惩罚用于特征学习的可行性并部分分析其效果。

原文摘要 · Abstract (English)

We study ridge-regularized log-density-ratio estimation in the Gaussian location model with a common covariance matrix. By affine invariance, the model is written as q $\sim$ N(0, I), p $\sim$ N($Δ$, I), with linear features, where $Δ$ is a mean vector. The variational estimator is the empirical Kullback-Leibler (KL) log-normalized fit with a squared L2-penalty on its nonconstant coefficient, and the spectral estimator recently introduced in [1] replaces a single variational problem by a continuum of ridge-regularized least-squares problems. We derive high-dimensional deterministic asymptotic equivalents when the numbers of observations and dimension tend to infinity with fixed ratios. The regularized variational limit is characterized by a scalar entropy minimization problem derived from the convex-Gaussian-min-max theorem (CGMT), while the regularized spectral limit follows from deterministic equivalents for resolvents of weighted sums of two independent Gaussian sample covariance matrices. We use these formulas to compare population risks, with experiments focused on fixed-signal aspect-ratio sweeps and optimized regularization. Our conclusion is that with many observations, under the criteria and asymptotic regimes analyzed here, the well-specified variational estimator has the smaller risk, while with fewer observations, the spectral estimator is favored because its covariance-based construction has lower variance. We also study how a nuclear penalty can be used and partially analyzed to perform feature learning.

密度比估计高维统计正则化谱方法

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。