测试提示工程对医疗大模型准确率与可信度的影响
Evaluating Prompt Engineering Techniques for Accuracy and Confidence Elicitation in Medical LLMs
- 对比不同提示风格和温度参数对模型表现的影响
- 思维链提示提升准确率但导致过度自信,需校准
- 小模型性能差,商业模型也缺乏可靠可信度
本文研究提示工程对应用于医疗场景的大语言模型在准确率和可信度获取方面的影响。基于涵盖多个专科的波斯语执业考试题的分层数据集,评估了五种模型——GPT-4o、o3-mini、Llama-3.3-70b、Llama-3.1-8b 和 DeepSeek-v3——在156种配置下的表现。这些配置包括温度设置(0.3、0.7、1.0)、提示风格(思维链、少样本、情感化、专家模仿)以及可信度量表(1-10、1-100)。使用AUC-ROC、Brier Score 和期望校准误差(ECE)评估可信度与实际表现的一致性。结果显示,思维链提示提升了准确率,但同时导致过度自信,凸显校准必要性;情感化提示进一步夸大可信度,可能引发错误决策。小型模型如 Llama-3.1-8b 在所有指标上表现较差,而专有模型虽准确率较高,但依然缺乏校准的可信度。结果表明,在高风险医疗任务中,提示工程必须兼顾准确性和不确定性表达。
原文摘要 · Abstract (English)
This paper investigates how prompt engineering techniques impact both accuracy and confidence elicitation in Large Language Models (LLMs) applied to medical contexts. Using a stratified dataset of Persian board exam questions across multiple specialties, we evaluated five LLMs - GPT-4o, o3-mini, Llama-3.3-70b, Llama-3.1-8b, and DeepSeek-v3 - across 156 configurations. These configurations varied in temperature settings (0.3, 0.7, 1.0), prompt styles (Chain-of-Thought, Few-Shot, Emotional, Expert Mimicry), and confidence scales (1-10, 1-100). We used AUC-ROC, Brier Score, and Expected Calibration Error (ECE) to evaluate alignment between confidence and actual performance. Chain-of-Thought prompts improved accuracy but also led to overconfidence, highlighting the need for calibration. Emotional prompting further inflated confidence, risking poor decisions. Smaller models like Llama-3.1-8b underperformed across all metrics, while proprietary models showed higher accuracy but still lacked calibrated confidence. These results suggest prompt engineering must address both accuracy and uncertainty to be effective in high-stakes medical tasks.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。