让大模型学会利用检查缺失信号判断病人预后,提升临床推理准确性。
Aligning Probabilistic Beliefs under Informative Missingness: LLM Steerability in Clinical Reasoning
- 通过提示工程引导模型理解检查项缺失的临床意义。
- 在真实重症监护数据上,模型对预后的概率判断更贴近真实分布。
- 适合关注医疗AI可解释性与临床决策支持的研究者。
大型语言模型(LLMs)在临床推理任务中应用日益广泛,这类任务需要基于现有证据生成校准的不确定性概率判断。然而,真实临床数据常存在缺失,且缺失模式往往蕴含患者预后的信息;例如,执行罕见检验可能反映医生的隐含怀疑。本文研究了是否可通过干预使LLM利用这种信息性缺失进行预后推断。为评估模型生成的概率信念与目标分布的对齐程度,我们分析了三种常见提示策略:显式序列化、指令引导和上下文学习。引入对对数损失的偏差-方差分解,以厘清性能提升的机制。在真实重症监护数据集上,发现显式结构引导和上下文学习可改善概率对齐,但模型若无精心设计的干预,无法自然利用信息性缺失。
原文摘要 · Abstract (English)
Large Language Models (LLMs) are increasingly deployed for clinical reasoning tasks, which inherently require eliciting calibrated probabilistic beliefs based on available evidence. However, real-world clinical data are frequently incomplete, with missingness patterns often informative of patient prognosis; for example, ordering a rare laboratory test reflects a clinician's latent suspicion. In this work, we investigate whether LLMs can be steered to leverage this informative missingness for prognostic inference. To evaluate how well LLMs align their verbalized probabilistic beliefs with an underlying target distribution, we analyze three common prompt-based interventions: explicit serialization, instruction steering, and in-context learning. We introduce a bias-variance decomposition of the log-loss to clarify the mechanisms driving gains in predictive performance. Using a real-world intensive care testbed, we find that while explicit structural steering and in-context learning can improve probabilistic alignment, the models do not natively leverage informative missingness without careful interventions.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。