提出新正则化方法,防止混合模型参数坍塌
Transcendental Regularization of Finite Mixtures:Theoretical Guarantees and Practical Limitations
- 用解析屏障函数惩罚似然,阻止成分坍缩
- 理论保证模型可识别、一致且鲁棒,实验稳定收敛
- 适合关注混合模型理论边界的研究者
有限混合模型广泛用于无监督学习,但通过EM算法进行最大似然估计时会出现成分坍缩的退化问题。本文提出超越正则化(transcendental regularization),一种具有解析屏障函数的惩罚似然框架,可在保持渐近效率的同时防止退化。由此产生的分布混合的超越算法(TAMD)具备强理论保证:可识别性、一致性与鲁棒性。实证上,TAMD有效稳定估计并防止坍缩,但在分类准确率上仅取得适度提升,揭示了高维无监督学习中混合模型的根本局限。本工作既提供新理论框架,也给出对实际局限的诚实评估,并开源实现于R包。
原文摘要 · Abstract (English)
Finite mixture models are widely used for unsupervised learning, but maximum likelihood estimation via EM suffers from degeneracy as components collapse. We introduce transcendental regularization, a penalized likelihood framework with analytic barrier functions that prevent degeneracy while maintaining asymptotic efficiency. The resulting Transcendental Algorithm for Mixtures of Distributions (TAMD) offers strong theoretical guarantees: identifiability, consistency, and robustness. Empirically, TAMD successfully stabilizes estimation and prevents collapse, yet achieves only modest improvements in classification accuracy-highlighting fundamental limits of mixture models for unsupervised learning in high dimensions. Our work provides both a novel theoretical framework and an honest assessment of practical limitations, implemented in an open-source R package.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。