构建可自动评估医疗AI临床能力的多维基准,更贴近真实诊疗场景。
GAPS: A Clinically Grounded, Automated Benchmark for Evaluating AI Clinicians
- 基于临床指南自动生成多维度评测题,覆盖认知深度、回答完整度等四方面。
- 90%题目与医生判断一致,模型在深层推理和安全性上表现明显不足。
- 适合医疗AI研发者、评估者使用,推动更安全可靠的临床应用。
当前医疗AI评估基准多依赖选择题或人工评分,难以反映真实临床所需的深度、鲁棒性和安全性。为此,我们提出GAPS框架,涵盖认知深度(Grounding)、回答完整度(Adequacy)、扰动鲁棒性(Perturbation)和安全性(Safety)。我们开发了完全自动化、基于指南的流水线,端到端构建符合GAPS标准的基准数据集:通过构建证据邻域,生成图与树双重结构表示,并自动生成跨G层级的问题。评分由一个模仿GRADE标准、基于PICO框架的DeepResearch代理在ReAct循环中合成评分标准,最终由大语言模型集成判别器打分。验证显示自动化问题质量高,与临床专家判断高度一致(90%一致性,Cohen's Kappa 0.77)。对先进模型的评估发现:性能随推理深度增加显著下降(G轴),回答完整性差(A轴),对对抗扰动极度敏感(P轴),且存在严重安全风险(S轴)。该自动化、临床锚定的方法为医疗AI系统提供了可复现、可扩展的严格评估路径。数据集GAPS-NSCLC-preview及评估代码已公开于https://huggingface.co/datasets/AQ-MedAI/GAPS-NSCLC-preview和https://github.com/AQ-MedAI/MedicalAiBenchEval。
原文摘要 · Abstract (English)
Current benchmarks for AI clinician systems, often based on multiple-choice exams or manual rubrics, fail to capture the depth, robustness, and safety required for real-world clinical practice. To address this, we introduce the GAPS framework, a multidimensional paradigm for evaluating Grounding (cognitive depth), Adequacy (answer completeness), Perturbation (robustness), and Safety. Critically, we developed a fully automated, guideline-anchored pipeline to construct a GAPS-aligned benchmark end-to-end, overcoming the scalability and subjectivity limitations of prior work. Our pipeline assembles an evidence neighborhood, creates dual graph and tree representations, and automatically generates questions across G-levels. Rubrics are synthesized by a DeepResearch agent that mimics GRADE-consistent, PICO-driven evidence review in a ReAct loop. Scoring is performed by an ensemble of large language model (LLM) judges. Validation confirmed our automated questions are high-quality and align with clinician judgment (90% agreement, Cohen's Kappa 0.77). Evaluating state-of-the-art models on the benchmark revealed key failure modes: performance degrades sharply with increased reasoning depth (G-axis), models struggle with answer completeness (A-axis), and they are highly vulnerable to adversarial perturbations (P-axis) as well as certain safety issues (S-axis). This automated, clinically-grounded approach provides a reproducible and scalable method for rigorously evaluating AI clinician systems and guiding their development toward safer, more reliable clinical practice. The benchmark dataset GAPS-NSCLC-preview and evaluation code are publicly available at https://huggingface.co/datasets/AQ-MedAI/GAPS-NSCLC-preview and https://github.com/AQ-MedAI/MedicalAiBenchEval.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。