用概率模型解释核范数正则化,提升低秩矩阵推断的准确性和效率。
A Probabilistic Basis for Low-Rank Matrix Learning

- 基于核范数构造概率分布,利用微分几何分析其性质
- 设计新MCMC算法,自动学习正则化参数λ,无需调参
- 在矩阵去噪与补全任务中显著提升精度与速度
低秩矩阵推断通常通过优化核范数正则化的代价函数实现。然而,尽管有众多计算方法,其背后的概率分布却缺乏深入理解。本文研究密度为 $f(X)/propto e^{-λig\ orm{X}\_\ast}$ 的概率分布,发现其基本属性可通过微分几何解析求解。利用这些性质,我们设计了改进的马尔可夫链蒙特卡洛(MCMC)算法用于低秩贝叶斯推断,并能自动学习正则化参数 $λ$,避免了在难以或无法调参时的人工调优。最后,该方法在数值实验中显著提升了低秩贝叶斯矩阵去噪与补全算法的准确性和效率。
原文摘要 · Abstract (English)
Low rank inference on matrices is widely conducted by optimizing a cost function augmented with a penalty proportional to the nuclear norm $\Vert \cdot \Vert_*$. However, despite the assortment of computational methods for such problems, there is a surprising lack of understanding of the underlying probability distributions being referred to. In this article, we study the distribution with density $f(X)\propto e^{-λ\Vert X\Vert_*}$, finding many of its fundamental attributes to be analytically tractable via differential geometry. We use these facts to design an improved MCMC algorithm for low rank Bayesian inference as well as to learn the penalty parameter $λ$, obviating the need for hyperparameter tuning when this is difficult or impossible. Finally, we deploy these to improve the accuracy and efficiency of low rank Bayesian matrix denoising and completion algorithms in numerical experiments.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。