arXiv:2503.05746cs.CYcs.LG2025-03

用无监督聚类方法实现95.31%自闭症筛查准确率

Unsupervised Clustering Approaches for Autism Screening: Achieving 95.31% Accuracy with a Gaussian Mixture Model

  • 采用高斯混合模型对无标签数据进行聚类分析
  • 在704人数据集上达到95.31%的分类映射准确率
  • 适合缺乏标注数据的自闭症早期筛查场景

自闭症谱系障碍(ASD)的诊断仍面临挑战,传统方法依赖有标签数据,获取成本高。本文探索四种无监督聚类算法——K均值、高斯混合模型(GMM)、层次聚类和DBSCAN——在公开的704名成人自闭症筛查数据集上的表现。通过交叉验证进行超参数调优后,高斯混合模型在映射原始自闭症/非自闭症标签时达到最高准确率95.31%。其他评估指标如调整兰德指数(ARI)和轮廓系数也显示了聚类内部一致性。数据经清洗、类别特征编码与标准化处理,并通过严格交叉验证比较各方法性能。结果表明,无监督方法在标注数据稀缺或昂贵场景下具有显著潜力,可助力早期筛查与高风险人群资源分配。

原文摘要 · Abstract (English)

Autism spectrum disorder (ASD) remains a challenging condition to diagnose effectively and promptly, despite global efforts in public health, clinical screening, and scientific research. Traditional diagnostic methods, primarily reliant on supervised learning approaches, presuppose the availability of labeled data, which can be both time-consuming and resource-intensive to obtain. Unsupervised learning, in contrast, offers a means of gaining insights from unlabeled datasets in a manner that can expedite or support the diagnostic process. This paper explores the use of four distinct unsupervised clustering algorithms K-Means, Gaussian Mixture Model (GMM), Agglomerative Clustering, and DBSCAN to analyze a publicly available dataset of 704 adult individuals screened for ASD. After extensive hyperparameter tuning via cross-validation, the study documents how the Gaussian Mixture Model achieved the highest clustering-to-label accuracy (95.31%) when mapped to the original ASD/NO classification (4). Other key performance metrics included the Adjusted Rand Index (ARI) and silhouette scores, which further illustrated the internal coherence of each cluster. The dataset underwent preprocessing procedures including data cleaning, label encoding of categorical features, and standard scaling, followed by a thorough cross-validation approach to assess and compare the four clustering methods (5). These results highlight the significant potential of unsupervised methods in assisting ASD screening, especially in contexts where labeled data may be sparse, uncertain, or prohibitively expensive to obtain. With continued methodological refinements, unsupervised approaches hold promise for augmenting early detection initiatives and guiding resource allocation to individuals at high risk.

自闭症筛查无监督学习聚类分析高斯混合模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。