arXiv:2606.17115cs.LGcs.AI2026-06

评估大模型在癌症多模态分析中的表现与可信度,发现融合模态可提升诊断准确性。

Probing, Fusion, and Trustworthiness: A Systematic Evaluation of Foundation Model Representations for Multimodal Cancer Analysis

论文配图:Probing, Fusion, and Trustworthiness: A Systematic Evaluation of Foundation Model Representations for Multimodal Cancer Analysis
图 1 · 摘自论文原文
  • 对比五种大模型在病理图像与组学数据上的单模态表现
  • 多模态融合在无主导模态时效果更优,且在分布外数据上表现稳定
  • 通过置信区间预测验证模型不确定性,提升临床决策可靠性

基础模型(FMs)已成为医学数据的强大表征提取工具,但其在分布偏移数据下的泛化能力仍待深入研究。本工作系统评估了五种基础模型在两个真实世界商业队列IH-BC和IH-NSCLC上的表现,涵盖全切片图像与转录组数据。在八个下游分类任务中,先对单模态探针性能进行基准测试,发现图像与组学表征携带互补预测信号。随后比较三种基于配对表征的图像-组学融合策略,评估多模态融合是否带来额外增益。进一步通过分位数回归校准评估所选单模态与多模态流程的可信度。结果表明,基础模型在分布外数据上表现具有竞争力,多模态融合仅在单一模态未主导信号时显著提升性能。置信区间预测显示,多数点预测失败案例中,真实诊断仍可被包含于预测集合内,凸显不确定性感知推理在临床支持中的价值。

原文摘要 · Abstract (English)

Foundation models (FMs) have emerged as powerful representation extractors for medical data, yet their generalizability to datasets under distribution shift remains underexplored. This work systematically evaluates FM-based representations on a suite of computational pathology tasks across two real-world commercial cohorts, IH-BC and IH-NSCLC, drawn from the licensed in-house (IH) oncology dataset. The analysis focuses on two modalities, whole-slide images and transcriptomic profiles, drawn from the IH multimodal data. We first benchmark unimodal probing performance across five FMs on eight downstream classification tasks, and find that image and omics representations carry complementary predictive signals. Then we investigate whether multimodal fusion can yield additional gains over unimodal baselines by comparing three image-omics fusion strategies built on paired representations. The trustworthiness of selected unimodal and multimodal pipelines is further assessed through conformal prediction. Our results show that FM representations achieve competitive performance on out-of-distribution data and that multimodal fusion helps mainly when no single modality dominates the signal. Conformal prediction reveals that in the majority of cases where a point prediction fails, the true diagnosis remains recoverable within the prediction set, reinforcing the value of uncertainty-aware inference for clinical support.

多模态癌症分析大模型可信推理

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。