提示工程并非提升医疗大模型表现的万能解药,效果因任务和模型而异。
Prompt engineering does not universally improve Large Language Model performance across clinical decision-making tasks
- 测试三款主流大模型在五类临床决策任务中的表现,发现诊断建议准确率高但检验推荐差。
- 提示工程仅改善了最低分任务(检验推荐),对其他任务反而有害。
- 精心匹配的示例未必优于随机示例,上下文多样性可能更重要。
大型语言模型(LLMs)在医学知识评估中展现潜力,但在真实临床决策中的实用性仍不明确。本研究评估了ChatGPT-4o、Gemini 1.5 Pro和LIama 3.3 70B在典型患者就诊全流程中的表现,涵盖36个病例的五个连续临床决策任务:鉴别诊断、紧急处理措施、相关检查、最终诊断和治疗建议。在两种温度设置(默认与零)下,所有模型在最终诊断上接近完美,但在相关检查推荐上表现较差,其余任务为中等水平。其中,ChatGPT在零温度下表现更优,而LIama在默认温度下更佳。进一步采用MedPrompt框架进行提示工程,结合目标式与随机动态少样本学习,结果表明:提示工程并非普适有效,虽显著提升了检查推荐任务表现,却对其他任务产生负面影响;目标式少样本提示未持续优于随机选择,说明紧密匹配示例的收益可能被上下文多样性损失抵消。结论指出,提示工程的效果高度依赖模型与任务,需采取定制化、情境感知的整合策略。
原文摘要 · Abstract (English)
Large Language Models (LLMs) have demonstrated promise in medical knowledge assessments, yet their practical utility in real-world clinical decision-making remains underexplored. In this study, we evaluated the performance of three state-of-the-art LLMs-ChatGPT-4o, Gemini 1.5 Pro, and LIama 3.3 70B-in clinical decision support across the entire clinical reasoning workflow of a typical patient encounter. Using 36 case studies, we first assessed LLM's out-of-the-box performance across five key sequential clinical decision-making tasks under two temperature settings (default vs. zero): differential diagnosis, essential immediate steps, relevant diagnostic testing, final diagnosis, and treatment recommendation. All models showed high variability by task, achieving near-perfect accuracy in final diagnosis, poor performance in relevant diagnostic testing, and moderate performance in remaining tasks. Furthermore, ChatGPT performed better under the zero temperature, whereas LIama showed stronger performance under the default temperature. Next, we assessed whether prompt engineering could enhance LLM performance by applying variations of the MedPrompt framework, incorporating targeted and random dynamic few-shot learning. The results demonstrate that prompt engineering is not a one-size-fit-all solution. While it significantly improved the performance on the task with lowest baseline accuracy (relevant diagnostic testing), it was counterproductive for others. Another key finding was that the targeted dynamic few-shot prompting did not consistently outperform random selection, indicating that the presumed benefits of closely matched examples may be counterbalanced by loss of broader contextual diversity. These findings suggest that the impact of prompt engineering is highly model and task-dependent, highlighting the need for tailored, context-aware strategies for integrating LLMs into healthcare.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。