arXiv:2605.11398cs.AIcs.CL2026-05

评测大模型识别医疗紧急程度的能力,发现模型常误判危急情况。

AcuityBench: Evaluating Clinical Acuity Identification and Uncertainty Alignment

论文配图:AcuityBench: Evaluating Clinical Acuity Identification and Uncertainty Alignment
图 1 · 摘自论文原文
  • 构建统一四级紧急度框架,整合5个真实医疗场景数据
  • 12个模型在明确病例中准确率差异大,对话模式易漏诊高危情况
  • 模型对模糊病例的判断过于确定,偏离医生真实不确定分布

我们提出AcuityBench,一个评估语言模型从用户医疗描述中识别恰当紧急程度的基准。现有健康基准多聚焦医学问答、广义健康交互或特定流程分诊,缺乏跨场景的紧急度识别统一评估。AcuityBench通过共享的四级紧急度框架(居家监测至紧急抢救),整合五类公开数据集:用户对话、在线论坛、临床案例、患者门户消息,共包含914例,其中697例为共识病例用于标准准确率评估,217例为医师确认的模糊病例用于不确定性评估。支持两种互补任务形式:问答式四分类,以及基于评分标准的自由对话响应评估。在12个前沿私有与开源模型中,我们发现明确病例的准确率和错误方向存在显著差异。对比任务形式显示:对话模式虽降低过度分诊,但增加低估风险,尤其在高紧急度情况下。在模糊病例中,无一模型接近医生判断分布,且模型预测集中度高于专家临床不确定性。我们还对高度模糊案例进行专家与模型判定比较,分析临床不确定性在标签分歧中的作用。结果表明,紧急度识别是独立的安全关键能力,AcuityBench可系统化比较并压力测试模型在真实医疗场景中引导用户到正确护理层级的能力。

原文摘要 · Abstract (English)

We introduce AcuityBench, a benchmark for evaluating whether language models identify the appropriate urgency of care from user medical presentations. Existing health benchmarks emphasize medical question answering, broad health interactions, or narrow workflow-specific triage tasks, but they do not offer a unified evaluation of acuity identification across these settings. AcuityBench addresses this gap by harmonizing five public datasets spanning user conversations, online forum posts, clinical vignettes, and patient portal messages under a shared four-level acuity framework ranging from home monitoring to immediate emergency care. The benchmark contains 914 cases, including 697 consensus cases for standard accuracy evaluation and 217 physician-confirmed ambiguous cases for uncertainty-aware evaluation. It supports two complementary task formats: explicit four-way classification in a QA setting, and free-form conversational responses evaluated with a rubric-based judge anchored to the same framework. Across 12 frontier proprietary and open-weight models, we find substantial variation in clear-case acuity accuracy and error direction. Comparing task formats reveals a systematic tradeoff: conversational responses reduce over-triage but increase under-triage relative to QA, especially in higher-acuity cases. In ambiguous cases, no model closely matches the distribution of physician judgments, and model predictions are more concentrated than expert clinical uncertainty. We also compare expert and model adjudication on a subset of maximally ambiguous cases, using those cases to examine the role of clinical uncertainty in label disagreement. Together, these results position acuity identification as a distinct safety-critical capability and show that AcuityBench enables systematic comparison and stress-testing of how well models guide users to the right level of care in real-world health use.

医疗AI紧急度评估模型可靠性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。