LLM诊断概率估计不准,需改进置信度评估方法
Position Paper On Diagnostic Uncertainty Estimation from Large Language Models: Next-Word Probability Is Not Pre-test Probability
- 用电子病历数据测试两个LLM的诊断概率估计
- 三种现有提取方法均存在明显偏差
- 临床决策支持者需关注模型置信度可靠性
大型语言模型(LLMs)正被探索用于诊断决策支持,但其在估计先验概率方面的能力仍有限,而先验概率对临床决策至关重要。本研究使用结构化电子健康记录数据,在三个诊断任务上评估了两种LLM(Mistral-7B 和 Llama3-70B),考察了三种当前提取LLM概率估计的方法,并揭示了它们的局限性。研究旨在强调提升LLM置信度估计技术的必要性。
原文摘要 · Abstract (English)
Large language models (LLMs) are being explored for diagnostic decision support, yet their ability to estimate pre-test probabilities, vital for clinical decision-making, remains limited. This study evaluates two LLMs, Mistral-7B and Llama3-70B, using structured electronic health record data on three diagnosis tasks. We examined three current methods of extracting LLM probability estimations and revealed their limitations. We aim to highlight the need for improved techniques in LLM confidence estimation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。