用分箱降低核密度估计计算成本,提升效率。
Binned semiparametric Bayesian networks for efficient kernel density estimation
- 通过数据分箱与稀疏张量,构建高效半参数贝叶斯网络
- 在保持精度的前提下,速度显著优于传统方法
- 适合高维数据的快速密度估计,尤其适合大规模数据
本文提出一种新型概率半参数模型,利用数据分箱降低非参数分布中核密度估计的计算开销。针对新模型,设计了两种新的条件概率分布:稀疏分箱核密度估计和傅里叶核密度估计。这两种分布通过使用稀疏张量并限制条件概率计算中的父节点数量,缓解了分箱模型常见的维度灾难问题。为验证该方法,我们进行了复杂度分析,并在合成数据及UCI机器学习仓库的数据集上开展多组对比实验,涵盖不同分箱规则、父节点限制、网格大小和样本数量,全面评估模型表现。结果表明,所提分箱半参数贝叶斯网络在结构学习和对数似然估计上与非分箱半参数贝叶斯网络无显著差异,但速度大幅提升。因此,该模型被证明是传统方法可靠且更高效的替代方案。
原文摘要 · Abstract (English)
This paper introduces a new type of probabilistic semiparametric model that takes advantage of data binning to reduce the computational cost of kernel density estimation in nonparametric distributions. Two new conditional probability distributions are developed for the new binned semiparametric Bayesian networks, the sparse binned kernel density estimation and the Fourier kernel density estimation. These two probability distributions address the curse of dimensionality, which typically impacts binned models, by using sparse tensors and restricting the number of parent nodes in conditional probability calculations. To evaluate the proposal, we perform a complexity analysis and conduct several comparative experiments using synthetic data and datasets from the UCI Machine Learning repository. The experiments include different binning rules, parent restrictions, grid sizes, and number of instances to get a holistic view of the model's behavior. As a result, our binned semiparametric Bayesian networks achieve structural learning and log-likelihood estimations with no statistically significant differences compared to the semiparametric Bayesian networks, but at a much higher speed. Thus, the new binned semiparametric Bayesian networks prove to be a reliable and more efficient alternative to their non-binned counterparts.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。