测试大模型在消化科题目的自信心,发现普遍过于自信。
Self-Reported Confidence of Large Language Models in Gastroenterology: Analysis of Commercial, Open-Source, and Quantized Models
- 用300道消化科考题测试多个大模型的自信心
- 顶尖模型自信心评分误差仅0.15-0.2,但仍有过度自信倾向
- 适合关注AI医疗可信度的研究者和临床AI应用开发者
本研究利用300道消化科执业考试风格的问题,评估了包括GPT、Claude、Llama、Phi、Mistral、Gemini、Gemma和Qwen在内的多个大语言模型的自报告响应确定性。表现最佳的模型(GPT-o1 preview、GPT-4o和Claude-3.5-Sonnet)的Brier分数为0.15-0.2,AUROC为0.6。尽管新模型性能有所提升,但所有模型均表现出一致的过度自信趋势。不确定性估计仍是大语言模型在医疗领域安全应用的重大挑战。
原文摘要 · Abstract (English)
This study evaluated self-reported response certainty across several large language models (GPT, Claude, Llama, Phi, Mistral, Gemini, Gemma, and Qwen) using 300 gastroenterology board-style questions. The highest-performing models (GPT-o1 preview, GPT-4o, and Claude-3.5-Sonnet) achieved Brier scores of 0.15-0.2 and AUROC of 0.6. Although newer models demonstrated improved performance, all exhibited a consistent tendency towards overconfidence. Uncertainty estimation presents a significant challenge to the safe use of LLMs in healthcare. Keywords: Large Language Models; Confidence Elicitation; Artificial Intelligence; Gastroenterology; Uncertainty Quantification
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。