用新机器学习方法提升脑肿瘤甲基化分类准确率。
A Novel Machine Learning Approach for Central Nervous System Tumor Classification from DNA Methylation

- 结合稀疏随机投影与多项逻辑回归,提升分类稳健性。
- 在2801样本队列中达96%准确率,在独立队列中达86%。
- 方法更可靠,对临床分型和治疗决策有实际影响。
DNA甲基化分析已成为中枢神经系统(CNS)肿瘤分类的强大工具,但跨队列迁移性、方法严谨性和多类别评估的鲁棒性仍面临挑战。本文提出一种新的、方法学严谨的机器学习方法,结合稀疏随机投影进行降维,再用多项逻辑回归进行分类。在广泛使用的参考分类器所设定的相同实验条件下评估,该方法在2,801样本参考队列中,经分层三重交叉验证,平均准确率达96%;在独立的1,104样本临床评估队列中,91类水平准确率为86%,甲基化类家族水平达93%。优于现有最先进方法的82%类水平一致性和88%家族水平一致性,分别提升约4和5个百分点。这一改进具有临床意义:诊断中分类准确率提高5个百分点可直接影响肿瘤亚型判定,进而影响治疗选择与后续临床决策。结果表明,基于更严格机器学习实践的模型,在多种评估场景下持续优于先前最优,显著提升了CNS肿瘤分类的可靠性。
原文摘要 · Abstract (English)
NA methylation profiling has become a powerful approach for central nervous system (CNS) tumor classification, yet important challenges remain regarding cross-cohort transferability, methodological correctness, and robust multiclass evaluation. In this work, we propose a novel and methodologically rigorous machine-learning approach for methylation-based CNS tumor classification that combines Sparse Random Projection for dimensionality reduction with multinomial logistic regression for classification. We evaluate the proposed approach in the same general experimental setting established by a widely used reference classifier. On the 2,801-sample reference cohort, our method achieves a mean accuracy of 96\% under stratified 3-fold cross-validation. On the independent 1,104-sample clinical evaluation cohort, it reaches 86\% accuracy at the 91-class level and 93\% when predictions are evaluated at the methylation class family level. These results improve upon the corresponding state-of-the-art reference figures of 82\% class-level concordance and 88\% family-level concordance, yielding absolute gains of approximately 4 and 5 percentage points, respectively. This improvement is clinically relevant: in a diagnostic setting, a 5-point increase in correct tumor classification can directly affect cancer subtype assignment and, in turn, influence treatment selection and downstream clinical decision-making. Our results show that the proposed model, grounded in stronger methodological practice in machine learning, consistently outperforms the previous state of the art across evaluation settings and can materially improve the reliability of CNS tumor classification.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。