无需交叉验证,自动估计概率张量的秩并提升建模精度。
Joint Bayesian Parameter and Model Order Estimation for Low-Rank Probability Mass Tensors
- 基于贝叶斯框架,联合推断概率张量的低秩成分与真实秩。
- 在合成与真实数据上实现更准的参数估计和自动秩选择。
- 适合需要高效、自适应建模的统计信号处理与推荐系统场景。
从观测数据中可靠估计一组随机变量的联合概率质量函数(PMF)是统计信号处理与机器学习的重要目标。将联合PMF建模为可进行低秩柯里-帕金森分解(CPD)的张量,已推动高效PMF估计算法的发展。然而,这些算法需预先指定张量的秩(模型阶数),而真实秩在实际应用中未知。通常通过观察验证误差或计算各类基于似然的信息准则从候选集选取合适秩,该过程可能耗费大量计算时间或硬件资源,或导致模型失配,影响准确性。本文提出一种新颖的贝叶斯框架,可从观测数据中同时估计低秩分量并推断其真实秩。我们构建贝叶斯PMF估计模型,为模型参数设置合适的先验分布,使秩可直接推断而无需交叉验证。随后,基于变分推断(VI)推导出确定性解,以近似各类模型参数的后验分布。数值实验涵盖合成数据及真实分类与物品推荐数据,验证了所提方法在估计精度、自动秩检测与计算效率方面的优势。
原文摘要 · Abstract (English)
Obtaining a reliable estimate of the joint probability mass function (PMF) of a set of random variables from observed data is a significant objective in statistical signal processing and machine learning. Modelling the joint PMF as a tensor that admits a low-rank canonical polyadic decomposition (CPD) has enabled the development of efficient PMF estimation algorithms. However, these algorithms require the rank (model order) of the tensor to be specified beforehand. In real-world applications, the true rank is unknown. Therefore, an appropriate rank is usually selected from a candidate set either by observing validation errors or by computing various likelihood-based information criteria, a procedure that could be costly in terms of computational time or hardware resources, or could result in mismatched models which affect the model accuracy. This paper presents a novel Bayesian framework for estimating the low-rank components of a joint PMF tensor and simultaneously inferring its rank from the observed data. We specify a Bayesian PMF estimation model and employ appropriate prior distributions for the model parameters, allowing the rank to be inferred without cross-validation.We then derive a deterministic solution based on variational inference (VI) to approximate the posterior distributions of various model parameters. Numerical experiments involving both synthetic data and real classification and item recommendation data illustrate the advantages of our VI-based method in terms of estimation accuracy, automatic rank detection, and computational efficiency.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。