为医疗AI设计真实场景下的评估基准,揭示高分模型在临床任务中的实际表现短板。
Measuring What Matters: Benchmarking Generative, Multimodal, and Agentic AI in Healthcare
- 构建覆盖文档、决策支持等多任务的医疗AI评估基准
- 实测前沿模型在临床任务中性能下降至0.53-0.85
- 适合关注AI临床落地可靠性的研究者与开发者
AI模型正日益部署于真实的临床环境,需在复杂且高风险的工作流程中保持可靠表现,而标准训练与验证数据集无法涵盖此类场景。评估此类系统需依赖基准测试:由任务、数据集和度量指标构成的结构化组合,以实现可复现、可比较的性能衡量。医疗AI的核心挑战并非单纯性能,而是缺乏系统性方法来衡量在真实条件下模型的可靠性、安全性与临床相关性。现有基准多聚焦于模型知识掌握程度,极少检验其在完整临床任务中的稳定表现。当前基准通过零散的数据集构建优化特定任务表现,导致前沿模型在医学执照考试中接近满分,但在真实临床任务中表现显著下滑:文档生成得分0.74–0.85,临床决策支持得分为0.61–0.76,行政与工作流任务仅得0.53–0.63。高分指标带来虚假的部署成熟错觉,且随着AI承担更关键临床角色,性能与实用性之间的差距持续扩大。若无严谨的基准设计框架,领域无法判断临床表现不佳是模型局限还是评估方式缺陷。
原文摘要 · Abstract (English)
AI models are increasingly deployed in live clinical environments where they must perform reliably across complex, high-stakes workflows that standard training and validation datasets were never designed to capture. Evaluating these systems requires benchmarks: structured combinations of tasks, datasets, and metrics that enable reproducible, comparable measurement of what a model can do. The central challenge in healthcare AI is not performance alone, but the absence of systematic methods to measure reliability, safety, and clinical relevance under real-world conditions. Most existing benchmarks test what a model knows; too few test whether it can perform reliably and without failing across the full complexity of real clinical tasks. Current benchmarks have accumulated through ad hoc dataset construction optimized for narrow task performance: frontier models achieve near-perfect scores on medical licensing examinations, but when evaluated across real clinical tasks, performance degrades sharply, scoring 0.74--0.85 on documentation, 0.61--0.76 on clinical decision support, and only 0.53--0.63 on administrative and workflow tasks \cite{medhelm}. High benchmark scores give a false sense of deployment readiness, and the gap between performance and utility widens precisely as AI systems take on more consequential clinical roles. Without a principled framework for benchmark design, the field cannot determine whether poor clinical performance reflects model limitations or failures in how performance is being measured.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。