测试大模型在用药安全中对患者信息变化的敏感性,发现多数模型无法正确响应。
Evaluating Counterfactual Sensitivity to Patient Information in Medication-Safety Reasoning

- 设计包含对照问题的医学推理基准,检验模型是否真用患者信息判断用药安全。
- 28个模型在反事实问题上准确率从63.6%降至45.1%,暴露决策盲区。
- 适合关注医疗AI可靠性、临床决策可信度的研究者和开发者。
当药物安全规则的患者特异性条件不满足时,仍应用该规则可能导致错误决策。现有医疗评估多依赖孤立固定场景,模型可能仅靠记忆药害关联而非基于患者信息判断规则适用性。为此,我们提出MedPIC-Bench,一个包含可溯源建议与专家验证问题的基准,融合遵循指南的问题与成对反事实问题——其中患者信息的可控变化决定规则是否适用。该基准包含467个标注了六个临床与推理维度的问题。在28个医学专用、通用及专有大模型中,所有模型在反事实问题上的表现均下降,平均准确率从63.6%降至45.1%。模型在患者属性直接提示已知禁忌时表现良好,但在需缩小或撤销安全警告时则困难重重。尽管模型理由常提及信息变化,但最终答案仍维持原判断。这一缺陷在医学专用模型中依然存在,其反事实表现平均落后于通用模型。MedPIC-Bench使条件规则的应用可测量,并揭示静态用药安全准确率无法反映患者特定可靠性。
原文摘要 · Abstract (English)
Applying a valid medication-safety rule when its patient-specific conditions are not met can produce an incorrect decision. Existing medical evaluations largely use isolated and fixed scenarios. A model may therefore answer correctly by recalling a drug-risk association without showing that it used patient information to decide whether the rule applies. To address this gap, we introduce MedPIC-Bench, a benchmark of source-verifiable recommendations and expert-validated questions for patient-specific medication-safety reasoning. It combines guideline-following questions with paired counterfactual questions in which a controlled change in patient information changes whether a rule applies. The benchmark contains 467 questions annotated along six clinical and reasoning dimensions. Across 28 medical-specific, general, and proprietary LLMs, every model performs worse on counterfactual questions, with mean accuracy falling from 63.6\% to 45.1\%. Models perform well when an explicit patient attribute directly signals a familiar contraindication, but struggle when patient information must narrow or withdraw a safety warning. Model rationales often acknowledge the changed patient information, yet the final answers retain the previous safety judgment. This vulnerability persists among medical-specific LLMs, whose average CF performance trails that of general LLMs. MedPIC-Bench therefore makes conditional rule application measurable and highlights the limitations of static medication-safety accuracy for assessing patient-specific reliability.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。