用异常病例重校医疗大模型信心,显著降低60%误判风险。
Enhancing Healthcare LLM Trust with Atypical Presentations Recalibration
- 利用罕见病例调整模型置信度估计,提升判断可靠性。
- 在三个医疗问答数据集上使校准误差降低约60%。
- 适合医疗决策、AI安全等高风险场景使用。
黑箱大语言模型在各类环境中日益广泛应用,尤其在高风险领域,其自信程度与不确定性表达至关重要。然而,这些模型常表现出过度自信,带来潜在风险与误判。现有校准方法主要针对通用推理数据集,效果有限。准确校准对决策支持和避免不良后果极为关键,但因任务复杂多样而困难重重。本文研究黑箱大模型在医疗场景中的校准偏差,提出一种新方法——异常表现重校准(Atypical Presentations Recalibration),通过利用异常临床表现来调节模型的置信度估计。该方法在三个医疗问答数据集上显著改善校准性能,校准误差减少约60%,优于原始口头置信度、思维链口头置信度等现有方法。同时,本文深入分析了异常性在重校准框架中的作用。
原文摘要 · Abstract (English)
Black-box large language models (LLMs) are increasingly deployed in various environments, making it essential for these models to effectively convey their confidence and uncertainty, especially in high-stakes settings. However, these models often exhibit overconfidence, leading to potential risks and misjudgments. Existing techniques for eliciting and calibrating LLM confidence have primarily focused on general reasoning datasets, yielding only modest improvements. Accurate calibration is crucial for informed decision-making and preventing adverse outcomes but remains challenging due to the complexity and variability of tasks these models perform. In this work, we investigate the miscalibration behavior of black-box LLMs within the healthcare setting. We propose a novel method, \textit{Atypical Presentations Recalibration}, which leverages atypical presentations to adjust the model's confidence estimates. Our approach significantly improves calibration, reducing calibration errors by approximately 60\% on three medical question answering datasets and outperforming existing methods such as vanilla verbalized confidence, CoT verbalized confidence and others. Additionally, we provide an in-depth analysis of the role of atypicality within the recalibration framework.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。