提出新数据增强算法AquaAugmentor,提升低维水质预测模型性能
AquaAugmentor: A Novel Feature Augmentation Algorithm for Water Potability Prediction

- 设计AquaAugmentor算法,通过特征空间增强改善水质分类数据
- 在多个模型上测试,显著提升准确率与AUC指标表现
- 适合关注环境监测与公共健康决策的机器学习研究者
安全饮水对健康、经济发展和可持续性至关重要。然而,由于水源数据的复杂性和可变性,准确分类水质仍是重大挑战。本文针对水体可饮用性预测问题,提出一种新型特征增强算法AquaAugmentor,旨在提升机器学习与深度学习模型在低维数据集上的预测性能。实验基于包含pH值、硬度、总固体、氯胺、硫酸盐等化学属性的水质数据集,评估了使用与不使用AquaAugmentor时各模型的分类表现,以测试准确率和AUC分数为评价标准。结果表明,该算法能有效增强模型性能,揭示了提升水质分类预测效果的关键技术路径。研究成果有助于推动安全饮水保障工作,为环境质量评估中的机器学习应用提供参考框架,支持科研人员、政策制定者及公共卫生官员基于可靠预测做出决策。
原文摘要 · Abstract (English)
Access to potable water is crucial for health, economic development, and sustainability. However, accurately classifying water quality remains a significant challenge due to the complexity and variability of water source data. This paper addresses the challenge of predicting water potability through machine learning and deep learning algorithms. It introduces a novel feature augmentation algorithm, AquaAugmentor, to enhance the predictive performance of these models for low-dimensional datasets. Utilizing a dataset that includes chemical attributes of water, such as pH, hardness, solids, chloramines, sulfate, and others. This study evaluates the performance of the models with and without AquaAugmentor. Each model applied to classify water as potable or non-potable and its performance is then evaluated and compared based on test accuracy and AUC score. The results highlight the strengths and limitations of our proposed algorithm, providing insights into the most effective techniques for improving the predictive performance of water quality classification. This study contributes to the broader efforts of ensuring safe water access and serves as a framework for employing machine learning in environmental quality assessments. The findings aim to assist researchers, policymakers, and public health officials in making informed decisions based on reliable machine learning predictions.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。