arXiv:2410.17263cs.LGcs.CY2024-10中稿 · ICLR被引 4

提出可解释偏见放大的理论,揭示模型设计如何加剧群体差异

An Effective Theory of Bias Amplification

  • 基于岭回归构建统一理论框架,分析正则化与数据分布对偏见的影响
  • 发现最优正则化强度可避免偏见放大,且参数增加未必改善少数群体性能
  • 理论预测与已有实证结果高度一致,适用于理解神经网络偏见机制

机器学习模型会捕获并放大数据中的偏见,导致不同社会群体间测试性能差异。为更好理解、评估和缓解此类偏见,亟需对模型设计选择和数据分布特性如何促成偏见有更深入的理论认识。本文在岭回归(含随机投影)背景下,建立了精确的解析理论,其中后者模拟了前馈神经网络在简化情形下的行为。该理论为机器学习偏见提供了统一且严格的解释,揭示了偏见放大、少数群体偏见等现象在多种特征与参数配置下的成因。例如,我们发现存在最优正则化惩罚或训练时长以避免偏见放大,且群体间测试误差差异可能不会随参数量增加而缓解。重要的是,理论预测与文献中关于机器学习偏见的实证观察高度吻合。我们在合成及半合成数据集上进行了广泛实验验证。

原文摘要 · Abstract (English)

Machine learning models can capture and amplify biases present in data, leading to disparate test performance across social groups. To better understand, evaluate, and mitigate these biases, a deeper theoretical understanding of how model design choices and data distribution properties contribute to bias is needed. In this work, we contribute a precise analytical theory in the context of ridge regression, both with and without random projections, where the former models feedforward neural networks in a simplified regime. Our theory offers a unified and rigorous explanation of machine learning bias, providing insights into phenomena such as bias amplification and minority-group bias in various feature and parameter regimes. For example, we observe that there may be an optimal regularization penalty or training time to avoid bias amplification, and there can be differences in test error between groups that are not alleviated with increased parameterization. Importantly, our theoretical predictions align with empirical observations reported in the literature on machine learning bias. We extensively empirically validate our theory on synthetic and semi-synthetic datasets.

偏见分析理论建模岭回归公平性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。