融合机器学习集成与大模型,提升心脏病预测准确率
Integrating Machine Learning Ensembles and Large Language Models for Heart Disease Prediction Using Voting Fusion
- 用集成学习+大模型投票融合,提升预测稳定性
- 混合模型达96.62%准确率,优于单一模型
- 适合医疗决策支持系统开发,尤其数据少时
心血管疾病是全球主要死因,亟需早期识别、精准风险分层与可靠的辅助决策技术。尽管机器学习(尤其是随机森林、XGBoost、LightGBM和CatBoost等集成方法)在建模复杂非线性患者数据方面表现优异,常超越逻辑回归,但大型语言模型(LLMs)提供了零样本和少量样本推理能力。本研究基于1,190例患者数据,比较传统机器学习模型(准确率95.78%,ROC-AUC 0.96)与开源大模型通过OpenRouter API的表现。结果表明,大模型单独使用时准确率为78.9%;而将机器学习集成与大模型推理结合,在Gemini 2.5 Flash下实现最佳性能(准确率96.62%,AUC 0.97)。这说明大模型在与机器学习协同时表现更佳,可显著增强不确定情况下的判断能力。集成学习仍是结构化表格预测的最优方案,但与大模型融合能带来微小提升,为更可靠的临床辅助工具开辟新路径。
原文摘要 · Abstract (English)
Cardiovascular disease is the primary cause of death globally, necessitating early identification, precise risk classification, and dependable decision-support technologies. The advent of large language models (LLMs) provides new zero-shot and few-shot reasoning capabilities, even though machine learning (ML) algorithms, especially ensemble approaches like Random Forest, XGBoost, LightGBM, and CatBoost, are excellent at modeling complex, non-linear patient data and routinely beat logistic regression. This research predicts cardiovascular disease using a merged dataset of 1,190 patient records, comparing traditional machine learning models (95.78% accuracy, ROC-AUC 0.96) with open-source large language models via OpenRouter APIs. Finally, a hybrid fusion of the ML ensemble and LLM reasoning under Gemini 2.5 Flash achieved the best results (96.62% accuracy, 0.97 AUC), showing that LLMs (78.9 % accuracy) work best when combined with ML models rather than used alone. Results show that ML ensembles achieved the highest performance (95.78% accuracy, ROC-AUC 0.96), while LLMs performed moderately in zero-shot (78.9%) and slightly better in few-shot (72.6%) settings. The proposed hybrid method enhanced the strength in uncertain situations, illustrating that ensemble ML is considered the best structured tabular prediction case, but it can be integrated with hybrid ML-LLM systems to provide a minor increase and open the way to more reliable clinical decision-support tools.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。