arXiv:2604.04239cs.LGcs.AI2026-04被引 1

多模态癌症生存模型预测概率严重不准,需警惕其临床误导风险。

Good Rankings, Wrong Probabilities: A Calibration Audit of Multimodal Cancer Survival Models

  • 首次系统审计多模态病理与基因融合模型的生存概率校准性
  • 3个模型在12/15折中校准失败,290次测试中166次显著偏差
  • 原生输出与后处理重构均存在校准问题,适合临床部署者必读

融合全切片图像与基因组数据的多模态深度学习模型在癌症生存预测上表现出较强的判别能力(以协和指数衡量)。然而,这些模型生成的生存概率是否准确校准仍缺乏研究。我们首次对多模态WSI-基因组生存架构进行了系统的逐折1-校准审计:实验A评估3个模型在TCGA-BRCA上的原生离散时间生存输出;实验B则评估11种架构在5种TCGA癌种中通过Breslow方法重构的生存曲线。实验A中,3个模型在多数折中未通过1-校准检验(15次测试中有12次在Benjamini-Hochberg校正后拒绝原假设)。在全部290次折级测试中,166次在中位事件时间处拒绝正确校准的原假设(FDR=0.05)。MCAT在GBMLGG上虽达到C-index 0.817,但所有5折均校准失败。基于门控的融合策略与更好校准相关,而双线性与拼接融合则不然。后处理Platt缩放可降低校准误差(如MCAT从5/5折失败降至2/5),且不影响判别性能。仅依赖协和指数不足以评估适用于临床的生存模型。

原文摘要 · Abstract (English)

Multimodal deep learning models that fuse whole-slide histopathology images with genomic data have achieved strong discriminative performance for cancer survival prediction, as measured by the concordance index. Yet whether the survival probabilities derived from these models - either directly from native outputs or via standard post-hoc reconstruction - are calibrated remains largely unexamined. We conduct, to our knowledge, the first systematic fold-level 1-calibration audit of multimodal WSI-genomics survival architectures, evaluating native discrete-time survival outputs (Experiment A: 3 models on TCGA-BRCA) and Breslow-reconstructed survival curves from scalar risk scores (Experiment B: 11 architectures across 5 TCGA cancer types). In Experiment A, all three models fail 1-calibration on a majority of folds (12 of 15 fold-level tests reject after Benjamini-Hochberg correction). Across the full 290 fold-level tests, 166 reject the null of correct calibration at the median event time after Benjamini-Hochberg correction (FDR = 0.05). MCAT achieves C-index 0.817 on GBMLGG yet fails 1-calibration on all five folds. Gating-based fusion is associated with better calibration; bilinear and concatenation fusion are not. Post-hoc Platt scaling reduces miscalibration at the evaluated horizon (e.g., MCAT: 5/5 folds failing to 2/5) without affecting discrimination. The concordance index alone is insufficient for evaluating survival models intended for clinical use.

生存分析模型校准多模态癌症预测

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。