arXiv:2606.31608cs.CL2026-06

用真人医生反馈评估大模型临床推理,发现其常因信息不足或表达冗长而误诊。

CLExEval: A Human-in-the-Loop Framework for Qualitative Evaluation of LLM Clinical Reasoning

论文配图:CLExEval: A Human-in-the-Loop Framework for Qualitative Evaluation of LLM Clinical Reasoning
图 1 · 摘自论文原文
  • 通过逐步隐藏信息并结合医生标注,检验大模型在真实临床场景下的推理能力。
  • 模型在信息减少时准确率从95%暴跌至32.5%,且68.6%正确推理未体现在最终诊断中。
  • 揭示大模型自评易失真,需真人验证才能可靠评估其临床可靠性。

大型语言模型(LLMs)在多个医学基准测试中表现优异,但其临床推理能力仍难以可靠评估。主要风险在于评估幻觉:流畅且结构良好的解释可能看似合理,实则诊断错误。我们提出CLExEval,一种基于真人医生参与的框架,用于在逐步信息屏蔽条件下评估LLM临床推理。该框架融合了5,600条专家医师标注与200条来自40个罕见病案例的推理轨迹。分析揭示三种常见失败模式:(i) 冗长偏差,GPT-4o-mini在信息稀缺下诊断准确率从95.0%降至32.5%;(ii) 隐性知识悖论,专用模型最高诊断潜力达92.5%却无法在冗长语境中稳定调用知识;(iii) 68.6%的推理与输出不一致,即正确诊断出现在推理过程但未反映在最终答案中。我们进一步在经人工验证的故障集(n=142)上评估了LLM作为裁判的范式:GPT-4o-mini批准了47.9%的临床错误输出,而HuatuoGPT-o1虽批准所有有效故障,但表现出正向自我偏好。结果表明,缺乏专家基准验证的自动化评估会严重高估临床可靠性。

原文摘要 · Abstract (English)

Large Language Models (LLMs) achieve strong results on many medical benchmarks, but their clinical reasoning remains difficult to evaluate reliably. A central risk is an evaluation illusion: fluent and well-structured explanations can appear clinically convincing even when the final diagnosis is incorrect. We introduce CLExEval, a human-in-the-loop framework for evaluating LLM clinical reasoning under progressive information masking. CLExEval combines 5,600 expert-physician annotations with 200 clinical reasoning traces derived from 40 rare diagnostic cases. Our analysis identifies three recurring failure patterns: (i) verbosity bias, where GPT-4o-mini's diagnostic accuracy drops from 95.0% to 32.5% under information scarcity; (ii) a hidden knowledge paradox, where a specialist model reaches 92.5% maximum diagnostic potential but fails to retrieve that knowledge reliably in verbose contexts; and (iii) a 68.6% reasoning-to-output mismatch, where correct diagnoses appear in reasoning traces but are not reflected in final answers. We further evaluate the LLM-as-a-Judge paradigm on a human-verified failure set (n = 142). GPT-4o-mini approved 47.9% of clinically incorrect outputs, while HuatuoGPT-o1 approved all validly scored failures and showed a positive self-preference bias. These results suggest that standalone automated clinical evaluations can substantially overestimate clinical reliability without expert-grounded validation.

临床推理大模型评估人类反馈

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。