提出诊断框架DECAT,判断多模态模型是否真懂生物机制。
When Are Multimodal Predictions Biologically Supported? A Diagnostic Evaluation Framework
- 用五项零参考指标+规则判定模型学的是共享生物学、单一模态或虚假关联
- 在8979名患者数据上验证,发现主流模型常误判无共享生物的场景
- 无需知道混杂因子就能发现隐藏偏差,适合医学多模态模型可信性评估
癌症多模态模型可实现高精度预测,但精度无法揭示其学习的是跨模态共享生物学、单一模态特异性信息,还是由混杂因子引发的虚假相关。本文提出DECAT——一种模型无关的后验评估框架,通过五项零参考指标与规则决策流程,将多模态表征分类为四种诊断情景。该框架基于学习到的表征运行,无需知晓具体混杂因子,证据不足时返回不确定结果。我们在合成数据(覆盖4类多模态模型,超2500个训练表征)和真实TCGA数据(8979名患者)上验证了DECAT,评估了多模态嵌入及五个预训练病理基础模型。结果显示,纠缠模型(如CLIP)在共享生物学检测上近乎完美,但在无共享生物学的情况下仍会错误声称存在共享,错误率随混杂强度增加而上升;更大的数据集和更强的表示反而导致更自信却错误的诊断。应用于无配对RNA的TCGA多模态嵌入及五个病理基础模型时,DECAT能检测出AUROC无法察觉的混杂,且无需混杂因子标签,经事后分层验证确认有效。
原文摘要 · Abstract (English)
Multimodal models in oncology can produce accurate predictions, but accurate prediction does not reveal whether the model has learned biology that is shared across modalities, biology confined to one modality, or spurious correlations that reflect confounders rather than genuine biology. We introduce DECAT, a model-agnostic post-hoc evaluation framework that classifies multimodal representations into four diagnostic scenarios for a given task and modality, using five null-referenced metrics and a rule-based decision procedure. The framework operates on learned representations, requires no knowledge of which specific confounder is present, and returns indeterminate when the evidence is insufficient. We validate DECAT on synthetic data across four multimodal model classes (over 2,500 trained representations) and on real data from 8,979 TCGA patients, evaluating both multimodal embeddings and five pretrained pathology foundation models. Entangled models (e.g., CLIP) achieve near-perfect shared biology detection but falsely claim shared biology in the majority of cases where it is absent on real foundation model embeddings. This false claim rate increases with confound strength so that larger cohorts and stronger representations produce more confident but still incorrect diagnoses. Applied to both multimodal TCGA embeddings and five pathology foundation models without paired RNA, DECAT detects confounding invisible to AUROC without requiring the confounder labels, as confirmed by post-hoc stratification.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。