用耦合模型融合临床与基因风险评分,发现双高分者预后最差。
Copula Based Fusion of Clinical and Genomic Machine Learning Risk Scores for Breast Cancer Risk Stratification
- 用高斯、Clayton等耦合模型建模临床与基因评分的联合分布
- 双高风险组5年癌症特异性死亡率显著更高,生存差异明显
- 方法可解释依赖关系,适合做多源数据融合研究
临床和基因表达模型可预测乳腺癌结局,但简单线性融合忽略了二者风险评分间的依赖关系。基于METABRIC数据,我们检验了建模临床与基因表达评分联合分布是否能提升5年癌症特异性死亡率的风险分层效果。定义临床与mRNA表达两个预测视图,通过5折交叉验证训练分类器并获取留出折叠概率。将得分转换为(0,1)^2上的伪观测值,拟合高斯、Clayton、Gumbel和Frank耦合模型。临床模型表现优于基因表达模型(AUC 0.783 vs 0.721)。Frank耦合具有最小拟合优度统计量,高斯表现相近。耦合融合未提升临床模型的ROC-AUC。然而,联合评分组间生存差异显著,双高分患者预后最差。竞争风险分析显示癌症死亡发生率同样呈现此模式。在独立TCGA数据中进行外部评估,使用共享预测因子与统一的5年总死亡终点,耦合融合、个体及简单融合评分的判别能力相当,置信区间重叠。所有模型均采用METABRIC校准。重复交叉验证置换重要性分析中无基因满足预设稳定性标准,故基因层面发现视为探索性。耦合模型提供风险评分间依赖关系的显式可解释描述,支持联合分组分析。本方法学研究未确立更优预测性能、已验证的临床风险分层或临床应用价值。
原文摘要 · Abstract (English)
Clinical and gene-expression models predict breast cancer outcomes, but simple linear fusion ignores dependence between their risk scores. Using METABRIC, we tested whether modeling the joint distribution of clinical and gene-expression scores improved stratification of 5-year cancer-specific mortality. We defined clinical and mRNA-expression predictor views, trained classifiers, and obtained out-of-fold probabilities through 5-fold cross-validation. The scores were transformed into pseudo-observations on (0,1)^2 and used to fit Gaussian, Clayton, Gumbel, and Frank copulas. The clinical model discriminated better than the gene-expression model (AUC 0.783 vs 0.721). Frank had the smallest goodness-of-fit statistic, with Gaussian performing similarly. Copula fusion did not improve ROC-AUC over the clinical model. However, joint score groups showed clear survival differences, with patients scoring high on both views having the poorest outcomes. Competing-risks analysis showed the same pattern for cancer-death incidence. We also conducted an external evaluation in independent TCGA data using shared predictors and a harmonized 5-year overall-mortality endpoint. Copula-fused, individual, and simple-fusion scores showed comparable discrimination with overlapping confidence intervals. All received the same METABRIC-based recalibration. No gene met the prespecified stability criterion under repeated cross-validated permutation importance, so gene-level findings were treated as exploratory. Copulas provide an explicit, interpretable description of dependence between clinical and gene-expression risk scores and support descriptive joint-group analyses. This methodological study does not establish superior prediction, validated clinical risk categories, or clinical utility.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。