梯度下降会放大数据偏差,让少数群体特征被忽视。
When majority rules, minority loses: bias amplification of gradient descent
- 构建多数-少数学习的理论框架,揭示训练机制偏向多数群体。
- 实验证明模型倾向学习多数群体特征,忽略少数群体独特信息。
- 适合关注公平性、数据偏差的研究者和工程师阅读。
尽管机器学习中存在偏见放大的大量实证证据,其理论基础仍不清晰。本文提出一个多数-少数学习任务的正式框架,揭示标准训练过程如何倾向于多数群体,产生忽略少数群体特性的刻板预测器。在假设总体分布与方差不平衡的前提下,分析发现三个关键结果:(i) 完整数据与刻板预测器之间的距离很近;(ii) 训练整个模型时,主要区域仅学习多数群体特征;(iii) 额外训练量存在下限。这些结论通过表格与图像分类任务中的深度学习实验得到验证。
原文摘要 · Abstract (English)
Despite growing empirical evidence of bias amplification in machine learning, its theoretical foundations remain poorly understood. We develop a formal framework for majority-minority learning tasks, showing how standard training can favor majority groups and produce stereotypical predictors that neglect minority-specific features. Assuming population and variance imbalance, our analysis reveals three key findings: (i) the close proximity between ``full-data'' and stereotypical predictors, (ii) the dominance of a region where training the entire model tends to merely learn the majority traits, and (iii) a lower bound on the additional training required. Our results are illustrated through experiments in deep learning for tabular and image classification tasks.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。