arXiv:2505.08775cs.CL2025-05被引 411

构建医疗大模型评估基准,真实对话测试模型表现与安全。

HealthBench: Evaluating Large Language Models Towards Improved Human Health

  • 基于5000通真实医疗对话,用262位医生制定的48562条标准评估。
  • GPT-4o得分32%,o3模型达60%,小模型成本降低25倍仍表现优异。
  • 提供共识版和难题版,适合研究医疗AI安全与实用性的团队使用。

我们提出HealthBench,一个开源基准,用于评估大语言模型在医疗领域的性能与安全性。该基准包含5000次模型与用户或医护人员之间的多轮对话,由262名医生制定对话专用评估标准。不同于以往的单选题或简答评测,HealthBench通过48,562个独特评估维度,在急诊、临床数据转换、全球健康等多重场景及准确性、指令遵循、沟通能力等行为维度上实现真实开放的评估。过去两年性能显示,初始进展平稳(如GPT-3.5 Turbo从16%提升至GPT-4o的32%),近期加速提升(o3模型达60%)。小型模型也显著进步:GPT-4.1 nano表现超越GPT-4o,成本仅为后者的1/25。我们还发布了两个变体:HealthBench Consensus(经医生共识验证的34个关键行为维度)和HealthBench Hard(当前最高分仅32%)。期望HealthBench推动更有利于人类健康的模型研发与应用。

原文摘要 · Abstract (English)

We present HealthBench, an open-source benchmark measuring the performance and safety of large language models in healthcare. HealthBench consists of 5,000 multi-turn conversations between a model and an individual user or healthcare professional. Responses are evaluated using conversation-specific rubrics created by 262 physicians. Unlike previous multiple-choice or short-answer benchmarks, HealthBench enables realistic, open-ended evaluation through 48,562 unique rubric criteria spanning several health contexts (e.g., emergencies, transforming clinical data, global health) and behavioral dimensions (e.g., accuracy, instruction following, communication). HealthBench performance over the last two years reflects steady initial progress (compare GPT-3.5 Turbo's 16% to GPT-4o's 32%) and more rapid recent improvements (o3 scores 60%). Smaller models have especially improved: GPT-4.1 nano outperforms GPT-4o and is 25 times cheaper. We additionally release two HealthBench variations: HealthBench Consensus, which includes 34 particularly important dimensions of model behavior validated via physician consensus, and HealthBench Hard, where the current top score is 32%. We hope that HealthBench grounds progress towards model development and applications that benefit human health.

医疗AI大模型评测基准测试LLM安全

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。