测试大模型在临床信息不全时能否判断能否下结论,发现它们常过早判断或过度回避。
ClinDet-Bench: Beyond Abstention, Evaluating Judgment Determinability of LLMs in Clinical Decision-Making
- 构建基于临床评分系统的基准,区分可判断和不可判断场景
- 模型在信息不全时既提前下结论又过度拒绝判断,正确率不足50%
- 适合评估医疗等高风险领域AI的安全性,推动更可靠的决策系统
临床决策常面临信息不全的挑战。临床专家需判断现有信息是否足以作出判断,过早下结论或无谓回避均可能危及患者安全。为评估大语言模型(LLMs)在此方面的能力,我们开发了ClinDet-Bench,一个基于临床评分系统的基准,将不完整信息情境分解为可判断与不可判断两类。识别可判断性需考虑所有关于缺失信息的假设(包括不太可能的情况),并验证结论在各类假设下是否成立。研究发现,近期的LLMs在信息不全时无法正确识别可判断性,既产生提前判断也出现过度回避,尽管它们能正确解释评分知识并在信息完整时表现良好。这表明现有基准不足以评估LLMs在临床环境中的安全性。ClinDet-Bench提供了一个评估可判断性识别能力的框架,有助于实现恰当的回避,具有在医学及其他高风险领域的潜在应用价值,且已公开可用。
原文摘要 · Abstract (English)
Clinical decisions are often required under incomplete information. Clinical experts must identify whether available information is sufficient for judgment, as both premature conclusion and unnecessary abstention can compromise patient safety. To evaluate this capability of large language models (LLMs), we developed ClinDet-Bench, a benchmark based on clinical scoring systems that decomposes incomplete-information scenarios into determinable and undeterminable conditions. Identifying determinability requires considering all hypotheses about missing information, including unlikely ones, and verifying whether the conclusion holds across them. We find that recent LLMs fail to identify determinability under incomplete information, producing both premature judgments and excessive abstention, despite correctly explaining the underlying scoring knowledge and performing well under complete information. These findings suggest that existing benchmarks are insufficient to evaluate the safety of LLMs in clinical settings. ClinDet-Bench provides a framework for evaluating determinability recognition, leading to appropriate abstention, with potential applicability to medicine and other high-stakes domains, and is publicly available.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。