用混合模型提升影像数据子群体聚类效果
Copula-based mixture model identification for subgroup clustering with imaging applications
- 基于核函数的混合模型允许各组分布形式灵活组合
- 在7万张MNIST图像和276例心脏影像上验证有效
- 适合需要灵活建模异质分布的医学影像分析
基于模型的聚类方法广泛应用于各类领域,但多数研究依赖于单一成分分布形式的典型混合模型,这一严格假设常难以满足。本文提出更灵活的基于核函数的混合模型(CBMMs),允许多样化的成分分布,通过灵活选择边缘分布和核函数形式实现。我们改进了广义迭代条件估计(GICE)算法,以无监督方式识别CBMMs,迭代估计边缘分布、核函数形式及其参数。该算法源自原用于切换马尔可夫模型识别的GICE,其设计考虑了实现时间的选择。所提的CBMM-GICE聚类方法在包含2000个样本的合成双聚类数据上进行了测试,并讨论了影响收敛性的因素。进一步与传统期望最大化(EM)算法识别的单一成分混合模型对比,在完整MNIST数据集(N=70,000)和真实心脏磁共振数据(N=276)上展示了其在影像分析中的价值。
原文摘要 · Abstract (English)
Model-based clustering techniques have been widely applied to various application areas, while most studies focus on canonical mixtures with unique component distribution form. However, this strict assumption is often hard to satisfy. In this paper, we consider the more flexible Copula-Based Mixture Models (CBMMs) for clustering, which allow heterogeneous component distributions composed by flexible choices of marginal and copula forms. More specifically, we propose an adaptation of the Generalized Iterative Conditional Estimation (GICE) algorithm to identify the CBMMs in an unsupervised manner, where the marginal and copula forms and their parameters are estimated iteratively. GICE is adapted from its original version developed for switching Markov model identification with the choice of realization time. Our CBMM-GICE clustering method is then tested on synthetic two-cluster data (N=2000 samples) with discussion of the factors impacting its convergence. Finally, it is compared to the Expectation Maximization identified mixture models with unique component form on the entire MNIST database (N=70000), and on real cardiac magnetic resonance data (N=276) to illustrate its value for imaging applications.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。