构建可控制可验证的医疗安全事件分类基准,提升大模型在临床风险判断中的可靠性。
PSEBench: A Controllable and Verifiable Benchmark for Evaluating LLMs in Patient Safety Event Triage

- 用条款卡片结构化法规文本,实现可审计的决策拆解。
- 生成5074个带真实答案的案例,支持信息补全与不确定情形测试。
- 适用于评估大模型在医疗安全领域的推理与拒答能力。
医疗安全事件分类是一项高风险任务,需判断临床事件是否符合特定地区政策要求,通常由安全专家人工完成。尽管大模型可能辅助此流程,但可靠评估受限于缺乏能捕捉基于证据的政策推理、主动补全不完整报告信息、以及在无法化解的模糊情况下合理拒答的基准。本文提出基于政策的构建方法,核心是“条款卡片”——一种将监管文本分解为可审计决策规范的结构化表示。结合锚点驱动实例化与闭环验证,该可扩展流程生成具备构造性真实答案的叙事,自然支持缺失信息补全和不确定变体生成。我们在明尼苏达州29项可报告不良健康事件上应用此方法,构建了包含5,074个案例的PSEBench基准,并配备智能体评估环境。对15个代表性大模型的评估揭示了稳定的能力趋势,验证了基准的有效性,并识别出实现可靠大模型医疗安全事件分类的关键改进方向。
原文摘要 · Abstract (English)
Patient safety event triage, determining whether a clinical event is reportable under jurisdiction-specific policy, is a high-stakes task typically performed manually by patient safety experts. Although LLMs may support this workflow, reliable evaluation is limited by the lack of benchmarks to capture evidence-grounded policy reasoning, proactive information seeking for incomplete reports, and principled abstention in irreducibly ambiguous cases. We address this gap with a policy-grounded construction methodology centered on the clause card, a structured representation that factorizes regulatory text into auditable decision specifications. Combining clause cards with anchor-driven instantiation and closed-loop verification, our scalable pipeline produces narratives with by-construction ground truth and naturally supports generating missing information and uncertain variants. We instantiate this method on Minnesota's 29 Reportable Adverse Health Events, producing PSEBench, a 5,074-case benchmark with an agentic evaluation environment. Evaluation on 15 representative LLMs reveals consistent capability trends, demonstrates the benchmark's utility, and identifies actionable gaps toward reliable LLM-based patient safety event triage.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。