LLM评估医疗回复完整性效果有限,与医生判断标准差异大。
Same Verdict, Different Reasons: LLM-as-a-Judge and Clinician Disagreement on Medical Chatbot Completeness

- 用三种评分粒度和三类模型测试LLM判别能力
- 准确率仅达随机水平(AUC 0.49–0.66),无法有效筛选不完整回复
- 医生与LLM结论一致时理由不同,不适合做临床自动筛查
LLM-as-a-Judge框架正被用于替代人工专家进行自动化评估,但在高风险医疗场景中的可靠性尚未验证。本文针对患者导向的医疗回复完整性检测,测试了三种评分粒度(General-Likert、Analytical-Rubric、Dynamic-Checklist)和三种基础模型,在两个由临床医生标注的数据集上展开评估,包括目前最大的公开医疗回复评价基准HealthBench。结果显示,LLM Judges对完整与不完整回复的区分能力仅处于接近随机水平(AUC 0.49–0.66);在需召回90%不完整回复的阈值下,临床医生仍需审查大部分数据,无显著分诊价值。即使模型与医生结论一致,其依据的理由也极少相同;分歧时,假阳性源于过度标记非关键缺失,而假阴性则反映根本性漏检。结果表明,LLM Judges与临床医生采用的根本不同的完整性标准,这削弱了其作为自主评估者或分诊过滤器在临床环境中的适用性。
原文摘要 · Abstract (English)
LLM-as-a-Judge frameworks are increasingly trusted to automate evaluation in place of human experts, yet their reliability in high-stakes medical contexts remains unproven. We stress-test this assumption for detecting incomplete patient-facing medical responses, evaluating three rubric granularities (General-Likert, Analytical-Rubric, Dynamic-Checklist) and three backbone models across two clinician-annotated datasets, including HealthBench, the largest publicly available benchmark for medical response evaluation. LLM Judges discriminate complete from incomplete responses at and slightly above near chance (AUC $0.49$--$0.66$); at the threshold required to recall $90\%$ of incomplete responses, clinicians must still review the vast majority of the dataset, offering no triage utility. Even when model and clinician verdicts agree, they rarely cite the same explanation; and when they diverge, false positives stem from over-flagging non-essential gaps while false negatives reflect outright detection failures. These results reveal that LLM Judges and clinicians apply fundamentally different completeness standards; a finding that undermines their use as autonomous evaluators or triage filters in clinical settings.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。