arXiv:2506.10769cs.CL2025-06中稿 · EACL 2026

首次系统评估临床问答中大模型的不确定性,发现不同专科和题型表现差异显著。

Mind the Gap: Benchmarking LLM Uncertainty and Calibration with Specialty-Aware Clinical QA and Reasoning-Based Behavioural Features

  • 基于推理行为特征设计轻量级新不确定性评估方法
  • 十种开源与专属模型在11个专科、6类问题上表现不一
  • 强调按专科和题型选型,避免盲目使用统一模型

在高风险临床问答场景中,大语言模型的不确定性量化至关重要。本文首次针对11个临床专科和6类问题类型,评估了十种开源模型(通用、生物医学、推理型)及代表性专有模型的不确定性估计效果。研究分析了基于得分的不确定性方法,提出一种基于推理模型行为特征的轻量级新方法,并考察了置信区间校准作为补充的集合预测策略。结果表明,不确定性可靠性并非单一属性,而是受临床专科与问题类型影响,存在校准与判别能力的差异。研究强调应根据临床应用场景选择或集成具备互补优势的模型。

原文摘要 · Abstract (English)

Reliable uncertainty quantification (UQ) is essential when employing large language models (LLMs) in high-risk domains such as clinical question answering (QA). In this work, we evaluate uncertainty estimation methods for clinical QA focusing, for the first time, on eleven clinical specialties and six question types, and across ten open-source LLMs (general-purpose, biomedical, and reasoning models), alongside representative proprietary models. We analyze score-based UQ methods, present a case study introducing a novel lightweight method based on behavioral features derived from reasoning-oriented models, and examine conformal prediction as a complementary set-based approach. Our findings reveal that uncertainty reliability is not a monolithic property, but one that depends on clinical specialty and question type due to shifts in calibration and discrimination. Our results highlight the need to select or ensemble models based on their distinct, complementary strengths and clinical use.

大模型临床问答不确定性模型评估

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。