建模数据差异性,让机器学习更可靠、公平、泛化更强
Data Heterogeneity Modeling for Trustworthy Machine Learning
- 把数据多样性纳入机器学习全流程设计
- 在医疗、金融等领域显著提升模型鲁棒性与公平性
- 适合关注模型可靠性与公平性的研究者参考
数据异质性在决定机器学习系统性能方面起着关键作用。传统算法通常以优化平均性能为目标,往往忽视数据集内部的固有差异,导致决策不可靠、跨领域泛化能力差、结果不公平以及错误科学推断等问题。因此,对数据异质性进行细致建模是构建可信数据驱动系统的关键。本文综述了异质性感知机器学习这一范式,系统性地将数据异质性考虑融入从数据收集、模型训练到评估与部署的整个机器学习流程。通过在医疗、农业、金融和推荐系统等关键领域应用,展示了该方法在提升模型鲁棒性、公平性和可靠性方面的显著优势,并有助于模型诊断与改进。此外,本文还探讨了未来方向,为数据挖掘领域提供了研究机遇,旨在推动异质性感知机器学习的发展。
原文摘要 · Abstract (English)
Data heterogeneity plays a pivotal role in determining the performance of machine learning (ML) systems. Traditional algorithms, which are typically designed to optimize average performance, often overlook the intrinsic diversity within datasets. This oversight can lead to a myriad of issues, including unreliable decision-making, inadequate generalization across different domains, unfair outcomes, and false scientific inferences. Hence, a nuanced approach to modeling data heterogeneity is essential for the development of dependable, data-driven systems. In this survey paper, we present a thorough exploration of heterogeneity-aware machine learning, a paradigm that systematically integrates considerations of data heterogeneity throughout the entire ML pipeline -- from data collection and model training to model evaluation and deployment. By applying this approach to a variety of critical fields, including healthcare, agriculture, finance, and recommendation systems, we demonstrate the substantial benefits and potential of heterogeneity-aware ML. These applications underscore how a deeper understanding of data diversity can enhance model robustness, fairness, and reliability and help model diagnosis and improvements. Moreover, we delve into future directions and provide research opportunities for the whole data mining community, aiming to promote the development of heterogeneity-aware ML.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。