首个面向全科临床决策的约束型大模型评测基准,揭示了高精度模型的安全短板。
GPAgentBench-2K: Benchmarking Large Language Model Agents in Complex Clinical Action Space

- 构建基于真实医生诊疗记录的约束马尔可夫决策过程环境,包含六类基础临床动作
- 16个前沿大模型在高风险病例中超过50%违反安全约束,暴露质量-安全差距
- 适合医疗AI安全评估、临床决策系统开发与约束强化学习研究者参考
大型语言模型(LLMs)在临床智能体应用中展现出巨大潜力,但现有评测将临床流程简化为静态预测或粗粒度动作空间的无约束马尔可夫决策过程(MDP)。为此,我们提出GPAgentBench-2K,首个针对全科临床决策的约束马尔可夫决策过程(CMDP)大模型评测基准,基于专家验证的真实全科医生诊疗记录构建。该环境涵盖六类基础临床动作,施加拓扑结构的工作流先验,并将安全导向的放弃决策作为首要结果。对16个先进大模型的评估显示,随着动作空间扩大,性能显著下降。关键发现是存在临床质量-安全差距:即使诊断准确率最高的前沿模型,在超过一半的高风险病例中仍违反安全约束。我们以约束组相对策略优化(C-GRPO)建立基线,表明显式建模约束虽优于无约束强化学习方法,但仍远未达到临床可接受的安全水平。
原文摘要 · Abstract (English)
Large Language Models (LLMs) show great potential as clinical agents, yet existing benchmarks reduce clinical workflows to static predictions or unconstrained Markov Decision Processes (MDPs) with coarse action sets. To address this, we introduce GPAgentBench-2K, the first Constrained MDP (CMDP) LLM-agent benchmark for primary-care clinical decision-making, constructed from expert-validated records of real-world GP encounters. Our environment models a full spectrum of six foundational clinical actions, imposes a topological workflow prior over the action space, and operationalizes safety-informed abstention as a first-class outcome. Evaluating 16 state-of-the-art LLMs reveals a significant performance degradation as the action space scales. Crucially, we uncover a clinical quality-safety gap: even frontier models with the highest diagnosis accuracy violate safety constraints in over half of high-risk cases. Finally, we establish a reference point using Constrained Group Relative Policy Optimization (C-GRPO), and show that while explicitly modeling constraints improves performance over unconstrained RL methods, it remains far from clinically acceptable safety.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。