通过反事实测试揭示医疗大模型是否真懂医学,而非仅依赖表面关联。
DeVisE: Behavioral Testing of Medical Large Language Models

- 设计反事实数据集,改变患者年龄、性别等变量测试模型反应。
- 发现不同模型对同一变量变化的响应差异大,但传统指标未暴露此问题。
- 适合评估医疗AI模型真实推理能力,尤其关注临床决策可靠性。
大型语言模型在临床决策支持中应用日益广泛,但现有评估很少揭示其输出是否反映真实的医学推理,还是仅依赖表面相关性。本文提出DeVisE(人口统计与生命体征评估)框架,通过可控的反事实实验探查细粒度临床理解能力。基于MIMIC-IV数据库中的重症监护室出院记录,构建包含单变量扰动(如年龄、性别、种族、生命体征)的真实与合成变体。在零样本设置下评估八种大模型(涵盖通用与医疗专用版本)。分析方法包括:(1) 输入级敏感性,衡量反事实扰动对困惑度的影响;(2) 下游推理表现,评估对预测住院时长和死亡率的影响。结果表明,标准任务指标掩盖了模型行为中的临床关键差异,各模型在响应反事实扰动的一致性和比例性上存在显著差异。
原文摘要 · Abstract (English)
Large language models (LLMs) are increasingly applied in clinical decision support, yet current evaluations rarely reveal whether their outputs reflect genuine medical reasoning or superficial correlations. We introduce DeVisE (Demographics and Vital signs Evaluation), a behavioral testing framework that probes fine-grained clinical understanding through controlled counterfactuals. Using intensive care unit (ICU) discharge notes from MIMIC-IV, we construct both raw (real-world) and template-based (synthetic) variants with single-variable perturbations in demographic (age, gender, ethnicity) and vital sign attributes. We evaluate eight LLMs, spanning general-purpose and medical variants, under zero-shot setting. Model behavior is analyzed through (1) input-level sensitivity, capturing how counterfactuals alter perplexity, and (2) downstream reasoning, measuring their effect on predicted ICU length-of-stay and mortality. Overall, our results show that standard task metrics obscure clinically relevant differences in model behavior, with models differing substantially in how consistently and proportionally they adjust predictions to counterfactual perturbations.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。