arXiv:2412.02189cs.AI2024-12被引 4

用出生后临床数据,机器学习可早期识别遗传病类型与亚型。

Comparative Performance of Machine Learning Algorithms for Early Genetic Disorder and Subclass Classification

  • 基于42项新生儿指标,用超参数调优的模型分类
  • 类别预测准确率77%,亚型识别最高达80%
  • 适合临床早期筛查,需更大数据验证

大量研究致力于发现特定遗传病,但其在广泛病种和亚型间的分类仍不明确。早期诊断有助于及时干预并改善预后。本研究利用出生或婴儿期可测得的基本临床指标,构建机器学习模型以实现生命早期的诊断。在包含22083个样本、42个特征(如家族史、新生儿指标、基础检验)的数据集上,实施了监督学习算法,并进行了广泛的超参数调优、特征工程与选择。开发了两种多分类器:一种用于预测疾病类别(线粒体、多因素、单基因),另一种用于亚型(共9种)。通过准确率、精确率、召回率及F1分数评估性能。CatBoost分类器在类别预测中达到最高准确率77%;SVM在亚型识别中最大准确率达80%。研究证明,利用基本临床数据进行机器学习建模,在多种遗传病的早期分类与诊断中具有可行性。结合基础临床指标的机器学习方法有望在验证后实现早期干预。未来需进一步研究以提升模型在该数据集上的表现。

原文摘要 · Abstract (English)

A great deal of effort has been devoted to discovering a particular genetic disorder, but its classification across a broad spectrum of disorder classes and types remains elusive. Early diagnosis of genetic disorders enables timely interventions and improves outcomes. This study implements machine learning models using basic clinical indicators measurable at birth or infancy to enable diagnosis in preliminary life stages. Supervised learning algorithms were implemented on a dataset of 22083 instances with 42 features like family history, newborn metrics, and basic lab tests. Extensive hyperparameter tuning, feature engineering, and selection were undertaken. Two multi-class classifiers were developed: one for predicting disorder classes (mitochondrial, multifactorial, and single-gene) and one for subtypes (9 disorders). Performance was evaluated using accuracy, precision, recall, and the F1-score. The CatBoost classifier achieved the highest accuracy of 77% for predicting genetic disorder classes. For subtypes, SVM attained a maximum accuracy of 80%. The study demonstrates the feasibility of using basic clinical data in machine learning models for early categorization and diagnosis across various genetic disorders. Applying ML with basic clinical indicators can enable timely interventions once validated on larger datasets. It is necessary to conduct further studies to improve model performance on this dataset.

遗传病机器学习早期诊断临床数据

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。