基于性能与复杂度评估,为医疗、通信、营销选最优机器学习模型
A Framework for Selection of Machine Learning Algorithms Based on Performance Metrices and Akaike Information Criteria in Healthcare, Telecommunication, and Marketing Sector
- 按数据属性和指标筛选算法,结合AIC控制模型复杂度
- 在8个真实数据集上验证,提升跨领域建模效率与准确率
- 适合需要自动化选型的医疗、电信、营销场景
互联网数据的爆炸式增长推动了人工智能、机器学习和深度学习在营销、电信和医疗领域的应用。本文聚焦于构建一个面向医疗、营销和电信三领域的机器学习算法选择框架。在医疗领域,解决心血管疾病预测(占全球死亡率28.1%)及胎儿健康状态分类问题,使用三个数据集;将算法分为急进型、懒惰型和混合型,依据数据特征、性能指标(准确率、精确率、召回率)及阿克伊克信息准则(AIC)评分进行筛选。实验采用来自三个领域的八个数据集进行验证。核心贡献是一个推荐框架,可根据输入属性自动识别最优模型,在性能与复杂度间取得平衡,提升多样化实际应用中的效率与准确性。该方法弥补了自动化模型选择的空白,具有跨学科部署的实际意义。
原文摘要 · Abstract (English)
The exponential growth of internet generated data has fueled advancements in artificial intelligence (AI), machine learning (ML), and deep learning (DL) for extracting actionable insights in marketing,telecom, and health sectors. This chapter explores ML applications across three domains namely healthcare, marketing, and telecommunications, with a primary focus on developing a framework for optimal ML algorithm selection. In healthcare, the framework addresses critical challenges such as cardiovascular disease prediction accounting for 28.1% of global deaths and fetal health classification into healthy or unhealthy states, utilizing three datasets. ML algorithms are categorized into eager, lazy, and hybrid learners, selected based on dataset attributes, performance metrics (accuracy, precision, recall), and Akaike Information Criterion (AIC) scores. For validation, eight datasets from the three sectors are employed in the experiments. The key contribution is a recommendation framework that identifies the best ML model according to input attributes, balancing performance evaluation and model complexity to enhance efficiency and accuracy in diverse real-world applications. This approach bridges gaps in automated model selection, offering practical implications for interdisciplinary ML deployment.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。