arXiv:2604.25415q-bio.NCcs.AI2026-04被引 2

15款前沿AI聊天机器人在精神科紧急分诊中,对危急情况识别准确但对非紧急情况过度预警。

One-shot emergency psychiatric triage across 15 frontier AI chatbots

  • 用112个真实对话案例测试AI对精神科紧急程度的判断能力。
  • 对危急情况识别准确率达94.3%,但中低风险误判率高达80.3%。
  • 适合关注AI医疗安全、精神健康筛查系统设计的研究者。

AI聊天机器人在健康咨询中应用日益广泛,但其在精神科分诊中的表现仍不明确。本研究评估了15款前沿AI聊天机器人在112个临床情景下的分诊表现,每个情景对应四个原始基准分诊等级(A:常规;B:一周内评估;C:24至48小时内评估;D:立即紧急处理)。情景涵盖9类精神症状和9项风险维度,共形成28个症状-风险组合,每组含4个不同等级的情景。每个情景以真实人类撰写的对话形式呈现,要求模型根据单条消息判断分诊等级。结果显示,在410次等级D测试中,有23次出现漏诊(5.6%),所有漏诊均被重新归为等级C。整体准确率在42.0%至71.8%之间,等级D最高(94.3%),等级B最低(19.7%)。平均有序误差为+0.47,表明存在总体高估风险倾向,且中间等级的判断波动最大。所有结果均经50名医生共识标签验证。当信息充分时,前沿AI可近乎零误差识别紧急情况,但在低至中等风险情境下表现出显著过度预警。

原文摘要 · Abstract (English)

AI chatbots are increasingly used for health advice, but their performance in psychiatric triage remains undercharacterized. Psychiatric triage is particularly challenging because urgency must often be inferred from thoughts, behavior, and context rather than from objective findings. We evaluated the performance of 15 frontier AI chatbots on psychiatric triage from realistic single-message disclosures using 112 clinical vignettes, each paired with 1 of 4 original benchmark triage labels: A, routine; B, assessment within 1 week; C, assessment within 24 to 48 hours; and D, emergency care now. Vignettes covered 9 psychiatric presentation clusters and 9 focal risk dimensions, organized into 28 presentation-by-risk groups. Each group contributed 4 distinct vignettes, with 1 vignette at each triage level. Each vignette was rendered as a realistic human-authored conversational query, and the AI chatbots were tasked with assigning a triage label from that disclosure. Emergency under-triage occurred in 23 of 410 level D trials (5.6%), and all under-triaged emergencies were reassigned to level C urgency. Across target models, average accuracy ranged from 42.0% to 71.8%. Accuracy was highest for level D vignettes (94.3%) and lowest for level B vignettes (19.7%). Mean signed ordinal error was positive (+0.47 triage levels), indicating net over-triage. Dispersion was highest around the middle triage levels. All results were confirmed relative to clinician consensus labels from 50 medical doctors. When presented with user messages containing sufficient clinical information, frontier AI chatbots thus recognized psychiatric emergencies as requiring urgent medical assessment with near-zero error rates, yet showed marked over-triage for low and intermediate risk presentations.

精神健康AI医疗分诊系统大模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。