arXiv:2607.01103cs.CL2026-07被引 1

LLM评估器看似达标,实则缺乏医生应有的谨慎与自我怀疑。

Clinician-Level Agreement Without Clinical Caution: LLM Evaluator Limits in Medical AI Benchmarking

论文配图:Clinician-Level Agreement Without Clinical Caution: LLM Evaluator Limits in Medical AI Benchmarking
图 1 · 摘自论文原文
  • 用3800个德语临床题构建首个标准化开放问答评测集。
  • 顶级模型与医生评分一致性达0.694,但无法体现对难题的回避。
  • 模型盲目给分,且偏好同架构模型,需警惕评估偏差。

开放式评估比选择题更具临床有效性,但存在评分瓶颈,促使使用大语言模型作为自动评价者。然而,这些评价者是否具备临床判断中的谨慎性尚不清楚。本文提出MedQADE,首个针对德语的标准化开放响应临床评测基准,包含3,800个由十名执业医生和九个大语言模型标注的题目。表现最佳的评估模型Gemini 3 Flash与医生评分一致性接近(κ=0.694 vs. κ=0.709),但置信区间较宽,解释受限。尽管统计上一致,自动化评估者几乎完全缺乏临床元认知:医生会随题目难度增加而选择不答,而前沿模型始终给出确定评分。此外,我们量化了系统性的谱系依赖偏差——模型更倾向于给同架构的模型高分,这一现象与语言无关。结果表明,统计一致性不代表临床谨慎性,评估独立性必须明确验证。

原文摘要 · Abstract (English)

Open-response evaluation provides stronger clinical validity than multiple-choice benchmarks but creates a scoring bottleneck that motivates automated LLM-asa-Judge approaches. Whether such evaluators replicate clinical calibration and caution, however, remains untested. We introduce MedQADE, the first standardised open-response clinical benchmark for German, a major clinical language lacking native evaluation infrastructure, comprising 3,800 items annotated by ten practising physicians and nine Large Language Model (LLM) evaluators. The top-performing evaluator model, Gemini 3 Flash, reached alignment consistent with the physician ceiling (\k{appa} = 0.694 vs. \k{appa} = 0.709), though wide confidence intervals limit interpretation. Despite this statistical alignment, automated evaluators exhibited near-absent clinical metacognition: physicians scaled abstention with item difficulty, while frontier models assigned definitive scores in every case. We additionally quantified systematic lineage-dependent biases, where models preferentially scored architectural siblings, an effect independent of language. These results show that statistical alignment does not ensure clinical caution, and that evaluator independence requires explicit verification.

医疗AI评测基准大模型评估临床谨慎

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。