提出混合集成特征选择方法,提升胰腺癌预后生物标志物发现的稳定性与可靠性。
Optimizing Prognostic Biomarker Discovery in Pancreatic Cancer Through Hybrid Ensemble Feature Selection and Multi-Omics Data
- 融合数据子采样与多种生存模型,结合嵌入式与包装式策略进行特征筛选。
- 通过帕累托前沿自动确定最优特征数,保持预测准确率同时实现模型稀疏。
- 在三个胰腺癌队列中识别出更少且更稳定的标志物,适合临床研究与高维生存分析。
利用高维多组学数据预测患者生存需要系统化的特征选择方法,以保证预测性能、特征稀疏性和结果可靠性。本文提出一种混合集成特征选择(hEFS)方法,结合数据子采样与多种生存预测模型,整合嵌入式与包装式策略。通过投票理论启发的聚合机制对组学特征进行排序,并采用帕累托前沿自动选择最优特征数量,无需人为设定阈值,平衡预测准确性与模型稀疏性。在三个胰腺癌队列的多组学数据上应用hEFS,相比传统的晚期融合CoxLasso模型,能识别出显著更少且更稳定的生物标志物,同时保持相当的判别性能。hEFS已集成于开源mlr3fselect R包,为高维生存建模和生物标志物发现提供稳健、可解释且具临床价值的工具。
原文摘要 · Abstract (English)
Prediction of patient survival using high-dimensional multi-omics data requires systematic feature selection methods that ensure predictive performance, sparsity, and reliability for prognostic biomarker discovery. We developed a hybrid ensemble feature selection (hEFS) approach that combines data subsampling with multiple prognostic models, integrating both embedded and wrapper-based strategies for survival prediction. Omics features are ranked using a voting-theory-inspired aggregation mechanism across models and subsamples, while the optimal number of features is selected via a Pareto front, balancing predictive accuracy and model sparsity without any user-defined thresholds. When applied to multi-omics datasets from three pancreatic cancer cohorts, hEFS identifies significantly fewer and more stable biomarkers compared to the conventional, late-fusion CoxLasso models, while maintaining comparable discrimination performance. Implemented within the open-source mlr3fselect R package, hEFS offers a robust, interpretable, and clinically valuable tool for prognostic modelling and biomarker discovery in high-dimensional survival settings.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。