用混合差异与矩匹配改进高维密度估计,更快更稳定。
Density estimation via mixture discrepancy and moments
- 用混合差异和矩比较替代星偏差,提升计算效率
- 高维下(30维)精度接近原方法,大样本提速2-20倍
- 适合需要快速高维密度估计的研究者使用
为将直方图统计推广至高维场景,已有基于星偏差的序列分割密度估计(DSP)方法,通过自适应二叉分割构建分段常数逼近。但星偏差计算复杂且不满足反射与旋转不变性。本文提出以混合差异和矩比较替代星偏差,分别得到基于混合差异的序列分割(DSP-mix)与基于矩的序列分割(MSP)。两者均具有计算可处理性及反射、旋转不变性。在重建贝塔混合、高斯混合与重尾柯西混合分布的实验中,最高达30维,结果表明:MSP在保持与原方法相当精度的同时,在大样本下速度提升2至20倍;DSP-mix在低维(d ≤ 6)测试中精度良好且效率提高,但在高维问题中因分割层级降低导致精度下降。
原文摘要 · Abstract (English)
With the aim of generalizing histogram statistics to higher dimensional cases, density estimation via discrepancy based sequential partition (DSP) has been proposed to learn an adaptive piecewise constant approximation defined on a binary sequential partition of the underlying domain, where the star discrepancy is adopted to measure the uniformity of particle distribution. However, the calculation of the star discrepancy is NP-hard and it does not satisfy the reflection invariance and rotation invariance either. To this end, we use the mixture discrepancy and the comparison of moments as a replacement of the star discrepancy, leading to the density estimation via mixture discrepancy based sequential partition (DSP-mix) and density estimation via moment-based sequential partition (MSP), respectively. Both DSP-mix and MSP are computationally tractable and exhibit the reflection and rotation invariance. Numerical experiments in reconstructing Beta mixtures, Gaussian mixtures and heavy-tailed Cauchy mixtures up to 30 dimension are conducted, demonstrating that MSP can maintain the same accuracy compared with DSP, while gaining an increase in speed by a factor of two to twenty for large sample size, and DSP-mix can achieve satisfactory accuracy and boost the efficiency in low-dimensional tests ($d \le 6$), but might lose accuracy in high-dimensional problems due to a reduction in partition level.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。