arXiv:2604.04175cs.LG2026-04被引 1

让医疗模型学会表达不确定,提升对缺失数据的鲁棒性。

Uncertainty-Aware Foundation Models for Clinical Data

  • 用概率分布代替单一嵌入,表示患者潜在状态的不确定性
  • 在多种临床任务中,预测准确率与不确定性校准均优于基线
  • 适合处理不完整、多模态医疗数据,尤其适用于缺失值场景

医疗基础模型通常借鉴自然语言和计算机视觉范式,强调大规模预训练与确定性表征,但临床数据本就稀疏、不规则且依赖模态。本文提出一种不确定性感知的基础建模框架,将每位患者表示为可能潜在状态的概率分布,而非单一嵌入。通过学习集合值表示并强制同一患者不同部分视图间的一致性,模型捕捉可推断的共性特征,同时显式编码认知不确定性。该框架融合多模态编码器与可扩展自监督目标,包括重构、对比对齐和分布正则化。在多种临床任务中,相比强基线,本方法提升了预测性能、缺失数据下的鲁棒性及不确定性校准能力。结果表明,建模‘未观测到的内容’而非仅关注‘已观测到的’,是医疗基础模型的关键归纳偏置。

原文摘要 · Abstract (English)

Healthcare foundation models have largely followed paradigms from natural language processing and computer vision, emphasizing large scale pretraining and deterministic representations over heterogeneous clinical data. However, clinical observations are inherently incomplete, reflecting sparse, irregular, and modality dependent measurements of an underlying physiologic state. In this work, we propose a framework for uncertainty aware foundation modeling that represents each patient not as a point embedding, but as a distribution over plausible latent states. By learning set valued representations and enforcing consistency across partial views of the same patient, the model captures what is invariantly inferable while explicitly encoding epistemic uncertainty. We integrate this formulation with multimodal encoders and scalable self supervised objectives, combining reconstruction, contrastive alignment, and distributional regularization. Across diverse clinical tasks, our approach improves predictive performance, robustness under missing data, and uncertainty calibration relative to strong baselines. These results suggest that modeling what is not observed rather than only what is constitutes a critical inductive bias for healthcare foundation models.

医疗AI不确定性建模基础模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。