arXiv:2607.16250cs.LGcs.AI2026-07

用多组学数据预测乳腺癌激素受体状态,随机森林表现最佳。

Benchmarking Machine Learning Models for Multi-Omics-Based Breast Cancer Prediction

论文配图:Benchmarking Machine Learning Models for Multi-Omics-Based Breast Cancer Prediction
图 1 · 摘自论文原文
  • 整合转录组、基因组和蛋白质组数据,用经典机器学习模型系统评估预测性能。
  • 多组学融合使准确率提升至90.3%,ROC-AUC达97.1%,优于单一数据模态。
  • 模型选出的ESR1等基因具有生物学意义,适合临床研究与精准医疗参考。

雌激素受体(ER)状态是乳腺癌诊断、预后和治疗选择的关键生物标志物。高通量测序技术的发展催生了多组学数据,为计算预测提供了互补分子信息。本研究基于TCGA-BRCA队列,系统评估了随机森林、XGBoost、LightGBM、CatBoost、支持向量机(SVM)和逻辑回归等经典机器学习模型在转录组(RNA表达)、基因组(拷贝数变异;CNV)和蛋白质组(RPPA)数据上的ER状态预测性能。实验采用分层训练-测试划分、分层五折交叉验证、类别不平衡处理及每折独立特征选择,确保评估可靠并避免数据泄露。结果表明,转录组数据提供最强预测信号,多组学整合带来稳定但有限的性能提升。在整合多组学设置下,随机森林表现最优,平衡准确率达90.3%,ROC-AUC为97.1%。模型反复选择出包括ESR1、PGR、FOXA1和GATA3在内的生物相关基因,验证了其生物学合理性。研究显示,经过严格正则化的经典机器学习方法对小规模高维基因组数据仍具高效性,多组学整合可为乳腺癌ER状态预测提供互补信息。

原文摘要 · Abstract (English)

Estrogen Receptor (ER) status is a critical biomarker in breast cancer diagnosis, prognosis, and treatment selection. Recent advances in high-throughput sequencing technologies have enabled the generation of multi-omics datasets that provide complementary molecular information for computational prediction tasks. This study presents a systematic benchmarking analysis of classical machine learning models for ER status prediction using transcriptomic (RNA expression), genomic (copy number variation; CNV), and proteomic (RPPA) data from the TCGA-BRCA cohort. A rigorous experimental framework incorporating stratified train-test splitting, stratified five-fold cross-validation, class imbalance handling, and fold-specific feature selection was employed to ensure reliable evaluation and prevent data leakage. Random Forest, XGBoost, LightGBM, CatBoost, Support Vector Machines (SVM), and Logistic Regression were evaluated across single-omic and multi-omic settings. Results demonstrated that RNA expression provided the strongest predictive signal, while multi-omic integration yielded modest but consistent improvements over individual modalities. Among all evaluated approaches, Random Forest achieved the best overall performance in the integrated multi-omic setting, obtaining a balanced accuracy of 90.3\% and an ROC-AUC of 97.1\%. Furthermore, recurrent selection of biologically relevant genes, including \textit{ESR1}, \textit{PGR}, \textit{FOXA1}, and \textit{GATA3}, supported the biological validity of the learned models. These findings indicate that carefully regularized classical machine learning methods remain highly effective for small, high-dimensional genomic datasets and that multi-omic integration provides complementary information for breast cancer ER status prediction.

乳腺癌多组学机器学习生物标志物

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。