arXiv:2604.10224cs.LGcs.AI2026-04

在AutoML中直接优化公平性,显著提升公平性且减少数据用量。

Exploring the impact of fairness-aware criteria in AutoML

  • 将多种公平性度量融入AutoML优化过程,覆盖从数据到模型的全流程。
  • 预测性能下降9.4%,但平均公平性提升14.5%,数据使用量减少35.7%。
  • 生成的模型更简洁,说明公平性无需以复杂模型为代价。

机器学习系统日益用于影响个体决策,但其常依赖有偏数据,导致对特定群体的不公平结果。随着自动化机器学习(AutoML)的普及,这一风险加剧,因多数框架仅关注提升预测性能。以往研究多仅在模型选择或超参数调优阶段引入公平性,忽略其他关键环节。本文研究将公平性直接集成至AutoML优化组件,构建完整机器学习流程(涵盖数据选择、转换、模型选择与调优)。针对公平性度量选择难题,采用互补的公平性指标以捕捉不同维度。相比仅追求性能的基线,该方法使预测性能下降9.4%,但平均公平性提升14.5%,数据使用量减少35.7%。同时,最终方案更简单完整,表明公平性不一定需要高复杂度模型实现。

原文摘要 · Abstract (English)

Machine Learning (ML) systems are increasingly used to support decision-making processes that affect individuals. However, these systems often rely on biased data, which can lead to unfair outcomes against specific groups. With the growing adoption of Automated Machine Learning (AutoML), the risk of intensifying discriminatory behaviours increases, as most frameworks primarily focus on model selection to maximise predictive performance. Previous research on fairness in AutoML had largely followed this trend, integrating fairness awareness only in the model selection or hyperparameter tuning, while neglecting other critical stages of the ML pipeline. This paper aims to study the impact of integrating fairness directly into the optimisation component of an AutoML framework that constructs complete ML pipelines, from data selection and transformations to model selection and tuning. As selecting appropriate fairness metrics remains a key challenge, our work incorporates complementary fairness metrics to capture different dimensions of fairness during the optimisation. Their integration within AutoML resulted in measurable differences compared to a baseline focused solely on predictive performance. Despite a 9.4% decrease in predictive power, the average fairness improved by 14.5%, accompanied by a 35.7% reduction in data usage. Furthermore, fairness integration produced complete yet simpler final solutions, suggesting that model complexity is not always required to achieve balanced and fair ML solutions.

AutoML公平性优化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。