测试大模型能否逻辑推断心梗风险,结果发现尚不靠谱。
Can Large Language Models Logically Predict Myocardial Infarction? Evaluation based on UK Biobank Cohort
- 用文本化风险因素让大模型评分,结合思维链检验推理逻辑。
- 模型预测准确率低于传统医学指标和机器学习模型。
- 提醒未来医疗大模型需懂医学数据与逻辑推理,不能仅靠语言生成。
背景:大型语言模型(LLMs)在临床决策支持中取得显著进展,但亟需高质量证据评估其基于真实世界医疗数据进行准确临床决策的能力与局限性。目的:量化评估当前最先进的大模型(ChatGPT 和 GPT-4)是否能通过逻辑推理预测心肌梗死(MI)发生风险,并对比不同模型表现以全面评估大模型性能。方法:本回顾性队列研究纳入英国生物样本库(UK Biobank)中2006至2010年招募的482,310名参与者,最终筛选出690名个体作为研究队列。将每位参与者的MI风险因素表格数据转化为标准化文本描述,供ChatGPT识别。通过提问要求其给出0到10分的风险评分,并采用思维链(Chain of Thought, CoT)策略评估其推理逻辑性。将ChatGPT的预测性能与已发表的医学评分工具、传统机器学习模型及其他大模型进行比较。结论:当前大模型尚未具备应用于临床医学的条件。未来医疗大模型应具备医学领域专业知识,能够理解自然语言与量化医疗数据,并实现逻辑推理。
原文摘要 · Abstract (English)
Background: Large language models (LLMs) have seen extraordinary advances with applications in clinical decision support. However, high-quality evidence is urgently needed on the potential and limitation of LLMs in providing accurate clinical decisions based on real-world medical data. Objective: To evaluate quantitatively whether universal state-of-the-art LLMs (ChatGPT and GPT-4) can predict the incidence risk of myocardial infarction (MI) with logical inference, and to further make comparison between various models to assess the performance of LLMs comprehensively. Methods: In this retrospective cohort study, 482,310 participants recruited from 2006 to 2010 were initially included in UK Biobank database and later on resampled into a final cohort of 690 participants. For each participant, tabular data of the risk factors of MI were transformed into standardized textual descriptions for ChatGPT recognition. Responses were generated by asking ChatGPT to select a score ranging from 0 to 10 representing the risk. Chain of Thought (CoT) questioning was used to evaluate whether LLMs make prediction logically. The predictive performance of ChatGPT was compared with published medical indices, traditional machine learning models and other large language models. Conclusions: Current LLMs are not ready to be applied in clinical medicine fields. Future medical LLMs are suggested to be expert in medical domain knowledge to understand both natural languages and quantified medical data, and further make logical inferences.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。