构建中文医疗大模型评估基准,精准衡量临床回应覆盖度。
ClinConsensus: A Physician-Calibrated Benchmark for Evaluating Clinical Rubric Coverage in Chinese Medical LLMs
- 设计36个专科、2500例专家标注病例,每例配30项医生制定评判标准。
- 11个前沿模型平均覆盖仅39.6%~52.1%,阈值达标率更低至17.8%~32.9%。
- 提出医生校准的覆盖度评分,适合医疗AI研发与临床应用评估者参考。
开放式医疗大模型评估仍缺乏基于医生共识的临床响应标准覆盖度支撑,尤其在本地化临床场景中更为薄弱。本文提出 extsc{ClinConsensus},一个涵盖36个专科、2500个专家标注病例的中文医疗评估基准,包含12类任务主题、多难度层级,以及面向公众与专业人员的不同语境。每个病例配有30项特定于案例的二元评判标准。为评估回答是否满足足够数量的医生制定标准,我们提出 extsc{Clinician-Anchored Coverage Score}(CACS),在阈值 $k=10$ 下进行实例化,并开发了结合GPT-5.1评分器与医生监督的Qwen3-8B判官的双评框架。对11个前沿大模型的评估显示:鲁棒性准确率在39.6%至52.1%之间,而CACS@10则在17.8%至32.9%之间,模型间存在19.2–21.9个百分点的显著差距。分层分析进一步揭示推理、证据使用、结构化提取、用药指导、随访建议和对话风格等方面存在显著差异。结果表明,医疗大模型评估应聚焦阈值化的、基于评判标准的临床覆盖度,而非平均部分正确性。
原文摘要 · Abstract (English)
Open-ended medical LLM evaluation remains weakly grounded in physician-calibrated coverage of clinically relevant response criteria, especially in localized clinical settings. We introduce \textsc{ClinConsensus}, a Chinese medical benchmark of 2{,}500 expert-curated cases spanning 36 specialties, 12 task themes, multiple difficulty levels, and lay-facing versus professional-facing settings. Each case is paired with 30 case-specific binary rubric criteria. To evaluate whether responses satisfy enough physician-authored criteria, we propose \emph{Clinician-Anchored Coverage Score} (CACS), a physician-calibrated threshold metric instantiated at \(k=10\), and develop a dual-judge framework combining a GPT-5.1 grader with a physician-supervised Qwen3-8B judge. Evaluating 11 frontier LLMs, we find a persistent coverage gap: Rubric Accuracy ranges from 39.6\% to 52.1\%, whereas CACS@10 ranges from 17.8\% to 32.9\%, leaving a 19.2--21.9 point gap across models. Stratified analyses further reveal substantial variation across reasoning, evidence use, structured extraction, medication instructions, follow-up, and dialogue register. These results suggest that medical LLM evaluation should measure thresholded, rubric-grounded clinical coverage rather than average partial correctness.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。