提出新指标衡量高斯混合模型学习难度,突破传统距离局限。
Pair Correlation Factor and the Sample Complexity of Gaussian Mixtures
- 引入配对相关因子(PCF)刻画分量均值聚类特性
- 在均匀球形情形下实现优于ε⁻²的样本复杂度上界
- 适用于需精准参数恢复的复杂混合模型研究
我们研究高斯混合模型(GMM)的学习问题,探讨哪些结构属性决定其样本复杂度。以往工作多将复杂度与各分量间的最小间距关联,但我们证明该视角不完整。本文提出配对相关因子(PCF),一个捕捉分量均值聚类特性的几何量。相较于最小间距,PCF更准确地反映参数恢复的难度。在均匀球形情况下,我们给出一种算法,其样本复杂度上界优于常规的ε⁻²,揭示了在某些情形下需要超过ε⁻²样本才能完成学习。
原文摘要 · Abstract (English)
We study the problem of learning Gaussian Mixture Models (GMMs) and ask: which structural properties govern their sample complexity? Prior work has largely tied this complexity to the minimum pairwise separation between components, but we demonstrate this view is incomplete. We introduce the \emph{Pair Correlation Factor} (PCF), a geometric quantity capturing the clustering of component means. Unlike the minimum gap, the PCF more accurately dictates the difficulty of parameter recovery. In the uniform spherical case, we give an algorithm with improved sample complexity bounds, showing when more than the usual $ε^{-2}$ samples are necessary.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。