arXiv:2607.25485cs.AIcs.CL2026-07

构建医疗健康智能体评估基准,测试其在真实场景下的临床安全与决策能力。

PatientAgentBench: A Benchmark Framework for Evaluating Patient-Facing Health AI Agents

论文配图:PatientAgentBench: A Benchmark Framework for Evaluating Patient-Facing Health AI Agents
图 1 · 摘自论文原文
  • 用模拟患者对话+工具调用的代理框架评估健康AI表现
  • 最强模型仅4.25/5分,三成模型三甲医院级诊断失误
  • 适合研究者、开发者和临床团队验证医疗智能体可靠性

健康AI正从问答系统演进为能与患者对话、分析病历并代行操作的智能体。现有评测多聚焦医学知识,以孤立问答为主。本文提出PatientAgentBench,一个面向患者端智能体的评估框架:将基础模型封装为代理,在包含医疗工具的沙箱中与模拟患者交互,由大模型作为评审员依据百余项临床标准对每段对话评分。为验证一致性,持证医生标注共享对话,结果显示评审与专家评分相邻一致率达79%-93%,媲美或超过医生间一致性。我们在1,200个相同场景下评测了10个模型(来自4个系列),发现分诊质量是关键差异维度——最弱模型通过率仅32%,最强达88%;弱模型常未做临床筛查即执行行政请求。临床安全与流程准确性也呈现相似趋势:最差模型频繁虚构无法执行的动作,而前沿模型仅在1%-3%情况下出错,主要源于未验证工具输出及紧急情况遗漏资源。更优模型虽缩小差距但未完全消除,最强者总分仅4.25/5。这些缺陷仅在持续使用工具的真实病历对话中显现,证明静态评测不足以应对自主性增强的医疗智能体。我们公开该可复现、经临床验证的评估框架,推动领域弥补此差距。

原文摘要 · Abstract (English)

Health AI is evolving from answering questions to agentic systems that converse with patients, reason about health records, and act on their behalf. Primary care guards against diagnostic errors and unsafe care; agents assisting in this domain warrant evaluation against the same risks. Current benchmarks focus on medical knowledge, assessed through isolated question-answering or clinician-facing tasks. PatientAgentBench benchmarks patient-facing agentic healthcare; it evaluates a foundation model, wrapped in an agent with a sandbox of healthcare tools, conversing with a simulated patient. Each conversation is scored by an LLM-as-a-Jury across six dimensions via over a hundred conversation-agnostic, clinician-grounded criteria. To validate alignment, licensed clinicians annotated shared conversations, yielding 79-93% adjacent agreement between jury and expert raters, on par with or exceeding clinician inter-rater agreement. We benchmarked 10 models across four families on the same 1,200 scenarios and found clinical gaps. Triage quality is the most discriminating dimension: pass rates rise from 32% for the weakest models to 88% for the strongest, with agents often acting on administrative requests without clinical screening. Clinical safety and workflow accuracy follow the same pattern: the weakest models fail often, fabricating unexecuted actions, while frontier models fail on only 1-3% of cases, from unverified tool outputs and omitted crisis resources in an emergency. More capable models narrow these gaps but do not close them; the strongest scores only 4.25 of 5 overall. These failures surface only in sustained, tool-using conversations against realistic patient records, confirming that static benchmarks are insufficient as healthcare agentic systems gain autonomy. We release the framework as a reproducible, clinician-validated evaluation standard to help the field close this gap.

医疗AI智能体评测临床安全基准测试

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。