arXiv:2503.16577cs.LGcs.AI2025-03被引 17

对比三种特征选择方法,发现互信息法在心脏病诊断中表现最佳。

Feature selection strategies for optimized heart disease diagnosis using ML and DL models

  • 用互信息、方差分析和卡方检验筛选临床特征,提升模型预测能力。
  • 互信息使神经网络准确率达82.3%,召回率达0.94,优于其他方法。
  • 简单模型如朴素贝叶斯用方差分析即可,适合资源有限场景。

心脏病仍是全球主要致病与致死原因,亟需高效诊断工具以实现早期发现和临床决策。本研究评估了互信息(MI)、方差分析(ANOVA)和卡方检验在11种机器学习(ML)与深度学习(DL)模型上的表现,基于临床指标数据集,使用精确率、召回率、AUC、F1分数和准确率等指标进行评价。结果显示,互信息在先进模型中表现最优,尤其在神经网络上达到最高准确率82.3%和召回率0.94;逻辑回归(准确率82.1%)与随机森林(准确率80.99%)也因使用互信息而性能提升。朴素贝叶斯与决策树在采用方差分析或卡方检验时分别获得76.45%和75.99%的准确率,具备计算效率优势。相反,KNN与支持向量机(SVM)无论使用何种特征选择方法,准确率均低于55%。该研究系统比较了特征选择策略,揭示其对模型性能的关键影响,为不同算法选择合适特征方法提供实用依据,助力心血管疾病诊断工具的优化。

原文摘要 · Abstract (English)

Heart disease remains one of the leading causes of morbidity and mortality worldwide, necessitating the development of effective diagnostic tools to enable early diagnosis and clinical decision-making. This study evaluates the impact of feature selection techniques Mutual Information (MI), Analysis of Variance (ANOVA), and Chi-Square on the predictive performance of various machine learning (ML) and deep learning (DL) models using a dataset of clinical indicators for heart disease. Eleven ML/DL models were assessed using metrics such as precision, recall, AUC score, F1-score, and accuracy. Results indicate that MI outperformed other methods, particularly for advanced models like neural networks, achieving the highest accuracy of 82.3% and recall score of 0.94. Logistic regression (accuracy 82.1%) and random forest (accuracy 80.99%) also demonstrated improved performance with MI. Simpler models such as Naive Bayes and decision trees achieved comparable results with ANOVA and Chi-Square, yielding accuracies of 76.45% and 75.99%, respectively, making them computationally efficient alternatives. Conversely, k Nearest Neighbors (KNN) and Support Vector Machines (SVM) exhibited lower performance, with accuracies ranging between 51.52% and 54.43%, regardless of the feature selection method. This study provides a comprehensive comparison of feature selection methods for heart disease prediction, demonstrating the critical role of feature selection in optimizing model performance. The results offer practical guidance for selecting appropriate feature selection techniques based on the chosen classification algorithm, contributing to the development of more accurate and efficient diagnostic tools for enhanced clinical decision-making in cardiology.

心脏病诊断特征选择机器学习医疗AI

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。