测试大模型在精神科诊断中何时该问、何时该停,避免瞎猜。
Ask Before You Diagnose: Safe-Psych, a Sequential Evaluation Benchmark for LLMs in Psychiatry

- 用逐步披露病历信息的方式,模拟真实临床判断过程。
- 超60%模型在信息不足时仍强行诊断,多数不主动求澄清。
- 适合关注医疗AI安全、模型决策可靠性的研究者使用。
大型语言模型(LLMs)在医疗决策支持中应用日益广泛,但临床证据常不完整或动态变化。当信息不足以做出可靠判断时,模型应请求澄清或放弃回答,而非给出无依据的结论。现有医学评测通常假设信息完备,无法评估模型在不确定性下的行为。我们提出 Safe-Psych,一个面向精神科临床诊断的序列化评测基准,包含超过1,000条真实精神科病历,按临床进展分段披露,每阶段由精神科医生标注行动:诊断(DIAGNOSE)、澄清(CLARIFY)或放弃(ABSTAIN)。我们在全信息与序列两种场景下评估多个先进LLM。结果表明,模型能力不等于判断校准:即使强模型在信息不全时也频繁错误诊断,约60%以上情况下未正确弃权;安全提示仅将错误从过早诊断转向过度弃权。在序列评测中,模型常在证据不足时就诊断,极少主动寻求澄清,除非被明确引导;提前诊断准确率显著低于及时诊断。总体显示,当前模型普遍缺乏识别临床信息缺失并请求补充的能力。我们公开 Safe-Psych,以推动医疗AI安全研究。
原文摘要 · Abstract (English)
Large language models (LLMs) are increasingly used for decision support in healthcare, but clinical evidence is often incomplete or evolving. When the available information is insufficient to support a reliable answer, models should request clarification or abstain rather than provide unsupported responses. Existing medical benchmarks, however, typically assume that complete information is available upfront. We introduce Safe-Psych, a sequential benchmark for evaluating how LLMs handle evolving diagnostic uncertainty in clinical psychiatry. Safe-Psych contains over 1,000 real-world psychiatric clinical notes segmented to simulate incremental evidence disclosure, with psychiatrist-derived action labels at each stage: DIAGNOSE, CLARIFY, or ABSTAIN. We evaluate multiple state-of-the-art LLMs in full-information and sequential settings. Our findings show that capability does not ensure calibration: even strong models struggle under incomplete clinical information, with under-abstention exceeding 60% for most models and safety-aware prompting reducing premature commitment only by shifting errors toward excessive abstention. In sequential evaluation, models frequently diagnose before sufficient evidence is available and rarely seek clarification unless explicitly prompted; these premature diagnoses are less accurate than on-time diagnoses. Overall, Safe-Psych reveals a limitation across the evaluated models: recognizing when clinical evidence is incomplete and additional information is needed. We release Safe-Psych to support research on improving LLM safety in healthcare.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。