大模型能用少量样本预测糖尿病,效果接近传统机器学习。
From Chat to Checkup: Can Large Language Models Assist in Diabetes Prediction?
- 用零次、一次、三次提示测试大模型在糖尿病预测中的表现。
- GPT-4o和Gemma-2-27B在少样本下准确率最高,超过传统模型。
- 开源模型表现较弱,需优化提示工程和领域微调。
尽管机器学习与深度学习广泛用于糖尿病预测,但大型语言模型(LLMs)在结构化数值数据上的应用仍不充分。本研究基于皮马印第安人糖尿病数据库(PIDD),测试了六种LLMs在零样本、单样本和三样本提示下的糖尿病预测效果,包括四种开源模型(Gemma-2-27B、Mistral-7B、Llama-3.1-8B、Llama-3.2-2B)和两种专有模型(GPT-4o、Gemini Flash 2.0)。同时与随机森林、逻辑回归和支持向量机(SVM)等传统模型对比。采用准确率、精确率、召回率和F1分数评估。结果显示,专有模型表现更优,其中GPT-4o与Gemma-2-27B在少样本设置中达到最高准确率;值得注意的是,Gemma-2-27B的F1分数超越传统模型。然而,不同提示策略间性能波动明显,且需领域微调。研究证明大模型可辅助医疗预测,建议未来聚焦提示工程与混合方法改进。
原文摘要 · Abstract (English)
While Machine Learning (ML) and Deep Learning (DL) models have been widely used for diabetes prediction, the use of Large Language Models (LLMs) for structured numerical data is still not well explored. In this study, we test the effectiveness of LLMs in predicting diabetes using zero-shot, one-shot, and three-shot prompting methods. We conduct an empirical analysis using the Pima Indian Diabetes Database (PIDD). We evaluate six LLMs, including four open-source models: Gemma-2-27B, Mistral-7B, Llama-3.1-8B, and Llama-3.2-2B. We also test two proprietary models: GPT-4o and Gemini Flash 2.0. In addition, we compare their performance with three traditional machine learning models: Random Forest, Logistic Regression, and Support Vector Machine (SVM). We use accuracy, precision, recall, and F1-score as evaluation metrics. Our results show that proprietary LLMs perform better than open-source ones, with GPT-4o and Gemma-2-27B achieving the highest accuracy in few-shot settings. Notably, Gemma-2-27B also outperforms the traditional ML models in terms of F1-score. However, there are still issues such as performance variation across prompting strategies and the need for domain-specific fine-tuning. This study shows that LLMs can be useful for medical prediction tasks and encourages future work on prompt engineering and hybrid approaches to improve healthcare predictions.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。