arXiv:2503.08870cs.LGstat.AP2025-03被引 1

在英国生物银行数据上对比八种机器学习模型,发现线性模型仍最高效可靠。

Comprehensive Benchmarking of Machine Learning Methods for Risk Prediction Modelling from Large-Scale Survival Data: A UK Biobank Study

  • 对比8种生存预测模型,涵盖从线性到深度学习的多种算法
  • 在样本量5000至25万间测试,发现惩罚型Cox模型表现最稳定
  • 强调线性模型在大规模数据中兼具高效与可扩展性,适合主流报告

预测建模对预防医学至关重要。尽管大规模前瞻性队列研究和多样化的机器学习(ML)算法推动了生存分析的发展,但选择最优算法仍具挑战。现有基准研究多基于小规模数据,其结论是否适用于整合组学与临床特征的大规模数据尚不明确。本研究在英国生物银行(UK Biobank, UKB)这一大规模前瞻性队列中,对八种不同的生存任务实现方法(涵盖线性与深度学习模型)进行基准测试,比较其在异质预测矩阵和不同终点下的区分能力与计算开销。此外,评估了模型在5,000至250,000例个体样本量下的可扩展性。结果表明,模型判别性能受终点频率和预测矩阵特性影响显著,其中(惩罚型)Cox比例风险模型表现极为稳健。值得注意的是,在观测数较大且预测矩阵较简单时,复杂模型更具优势。计算开销差异显著,本文还提出针对当前不可行实现的解决方案。结论指出,最优模型选择依赖于样本量、终点频率及预测矩阵特性,为类似数据集的研究者提供重要参考。同时,本文展示线性模型在大规模风险建模中仍具高度有效性与可扩展性,建议在报告非线性模型时一并呈现。

原文摘要 · Abstract (English)

Predictive modelling is vital to guide preventive efforts. Whilst large-scale prospective cohort studies and a diverse toolkit of available machine learning (ML) algorithms have facilitated such survival task efforts, choosing the best-performing algorithm remains challenging. Benchmarking studies to date focus on relatively small-scale datasets and it is unclear how well such findings translate to large datasets that combine omics and clinical features. We sought to benchmark eight distinct survival task implementations, ranging from linear to deep learning (DL) models, within the large-scale prospective cohort study UK Biobank (UKB). We compared discrimination and computational requirements across heterogenous predictor matrices and endpoints. Finally, we assessed how well different architectures scale with sample sizes ranging from n = 5,000 to n = 250,000 individuals. Our results show that discriminative performance across a multitude of metrices is dependent on endpoint frequency and predictor matrix properties, with very robust performance of (penalised) COX Proportional Hazards (COX-PH) models. Of note, there are certain scenarios which favour more complex frameworks, specifically if working with larger numbers of observations and relatively simple predictor matrices. The observed computational requirements were vastly different, and we provide solutions in cases where current implementations were impracticable. In conclusion, this work delineates how optimal model choice is dependent on a variety of factors, including sample size, endpoint frequency and predictor matrix properties, thus constituting an informative resource for researchers working on similar datasets. Furthermore, we showcase how linear models still display a highly effective and scalable platform to perform risk modelling at scale and suggest that those are reported alongside non-linear ML models.

生存分析机器学习风险预测大规模数据

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。