arXiv:2504.18686cs.LGcs.IT2025-04

用最小描述长度优化分箱,结合张量分解实现更平滑的多变量密度估计。

A Unified MDL-based Binning and Tensor Factorization Framework for PDF Estimation

  • 基于MDL与分位数切分自动确定最优分箱,避免均匀分箱的缺陷。
  • 在真实干豆分类数据集上准确率显著优于传统直方图方法。
  • 适合需要平滑导数的场景,如梯度优化与非参数判别分析。

可靠的密度估计在统计学和机器学习中至关重要。许多实际场景中,数据应建模为能捕捉复杂多峰模式的混合成分密度。然而,传统的基于均匀直方图的密度估计器难以捕捉局部变化,尤其当底层分布高度非均匀时。此外,直方图固有的不连续性给需要平滑导数的任务(如基于梯度的优化、聚类和非参数判别分析)带来挑战。本文提出一种新的非参数多变量概率密度函数(PDF)估计方法,采用基于最小描述长度(MDL)的分箱策略,结合分位数切割。该方法基于张量分解技术,利用联合概率张量的典型多线性分解(CPD)。我们在合成数据和一个具有挑战性的真实干豆分类数据集上验证了该方法的有效性。

原文摘要 · Abstract (English)

Reliable density estimation is fundamental for numerous applications in statistics and machine learning. In many practical scenarios, data are best modeled as mixtures of component densities that capture complex and multimodal patterns. However, conventional density estimators based on uniform histograms often fail to capture local variations, especially when the underlying distribution is highly nonuniform. Furthermore, the inherent discontinuity of histograms poses challenges for tasks requiring smooth derivatives, such as gradient-based optimization, clustering, and nonparametric discriminant analysis. In this work, we present a novel non-parametric approach for multivariate probability density function (PDF) estimation that utilizes minimum description length (MDL)-based binning with quantile cuts. Our approach builds upon tensor factorization techniques, leveraging the canonical polyadic decomposition (CPD) of a joint probability tensor. We demonstrate the effectiveness of our method on synthetic data and a challenging real dry bean classification dataset.

密度估计张量分解非参数多变量

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。