评测大模型在真实临床对话中的表现,助力可信医疗AI发展。
HealthBench Professional: Evaluating Large Language Models on Real Clinician Chats

- 基于真实医生与ChatGPT的对话构建评估基准
- GPT-5.4在临床任务中超越人类医生表现
- 包含对抗性测试样本,适合医疗AI研发者使用
数百万临床医生使用ChatGPT辅助临床工作,但针对其典型应用场景的评估仍有限。我们提出HealthBench Professional,一个面向真实临床任务的开放基准,涵盖诊疗咨询、文书撰写与医学研究三大核心场景。每个案例均来自医生撰写的对话,并由三名以上医生分三阶段评审打分。共从15,079个候选样本中筛选出高质量、具代表性且对当前OpenAI前沿模型具挑战性的样本,难度提升约3.5倍。约三分之一样本涉及医生主动进行对抗性测试。作为基线,我们收集了人类医生在不限时、专科匹配、可联网条件下的响应。最佳系统GPT-5.4 in ChatGPT for Clinicians显著优于基础版GPT-5.4、其他模型及人类医生。该基准旨在为医疗AI社区提供追踪模型进展的工具,推动可信赖临床系统建设。
原文摘要 · Abstract (English)
Millions of clinicians use ChatGPT to support clinical care, but evaluations of the most common use cases in model-clinician conversations are limited. We introduce HealthBench Professional, an open benchmark for evaluating large language models on real tasks that clinicians bring to ChatGPT in the course of their work. The benchmark is organized around three common use cases central to clinical practice: care consult, writing and documentation, and medical research. Each example includes a physician-authored conversation with ChatGPT for Clinicians and is scored via rubrics written and iteratively adjudicated by three or more physicians across three phases. HealthBench Professional examples were carefully selected for quality, representativeness, and difficulty for OpenAI's current frontier models, to enable continued measurement of progress. Difficult examples for recent OpenAI models were enriched by roughly 3.5 times relative to the candidate pool of 15,079 examples. Additionally, about one-third of examples involve physicians conducting deliberate adversarial testing of models. As a strong baseline, we also collected human physician responses for all tasks (unbounded time, specialist-matched, web access). The best scoring system, GPT-5.4 in ChatGPT for Clinicians, outperforms base GPT-5.4, all other models, and human physicians. We hope HealthBench Professional provides the healthcare AI community a measure to track frontier model progress in real-world clinical tasks and build systems that clinicians can trust to improve care.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。