arXiv:2409.10191cs.CL2024-09被引 2

GPT-4在谵妄风险预测中表现不佳,远不如专业医疗AI模型准确。

LLMs for clinical risk prediction

  • 对比GPT-4与clinalytix Medical AI在谵妄风险预测中的表现
  • GPT-4对阳性病例识别率低,概率估计不可靠,准确率显著低于专业模型
  • 提示大模型在复杂临床决策中需人工辅助,不宜独立使用

本研究比较了GPT-4与clinalytix Medical AI在预测谵妄发生临床风险方面的有效性。结果显示,GPT-4在识别阳性病例方面存在显著缺陷,难以提供可靠的谵妄风险概率估计,而clinalytix Medical AI展现出更优的准确性。对大型语言模型输出的深入分析揭示了这些差异的潜在原因,与现有文献报道的局限性一致。结果凸显了大模型在准确诊断及解析复杂临床数据方面面临的挑战。尽管大模型在医疗领域具有巨大潜力,但目前尚不适合独立用于临床决策。应将其作为辅助工具,配合临床专业知识使用。持续的人类监督对确保患者和医护人员获得最佳结果至关重要。

原文摘要 · Abstract (English)

This study compares the efficacy of GPT-4 and clinalytix Medical AI in predicting the clinical risk of delirium development. Findings indicate that GPT-4 exhibited significant deficiencies in identifying positive cases and struggled to provide reliable probability estimates for delirium risk, while clinalytix Medical AI demonstrated superior accuracy. A thorough analysis of the large language model's (LLM) outputs elucidated potential causes for these discrepancies, consistent with limitations reported in extant literature. These results underscore the challenges LLMs face in accurately diagnosing conditions and interpreting complex clinical data. While LLMs hold substantial potential in healthcare, they are currently unsuitable for independent clinical decision-making. Instead, they should be employed in assistive roles, complementing clinical expertise. Continued human oversight remains essential to ensure optimal outcomes for both patients and healthcare providers.

临床预测大模型医疗AI

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。