统一解释模型偏差的等效性,揭示公平性与鲁棒性的内在联系
When Are Learning Biases Equivalent? A Unifying Framework for Fairness, Robustness, and Distribution Shift
- 用信息论定义偏差为条件独立性破坏,建立理论等价关系
- 发现虚假相关强度α对应子群体失衡比r≈(1+α)/(1-α)
- 实验证明跨数据集和模型的偏差方法可迁移,误差小于3%
机器学习系统存在多种失效模式:对保护群体的不公平、对虚假相关性的脆弱性、少数子群体表现差,这些通常由不同研究社区独立探讨。本文提出一个统一的理论框架,刻画不同偏差机制在模型性能上产生定量等效效果的条件。通过信息论度量将偏差形式化为条件独立性的违反,证明了虚假相关、子群体偏移、类别不平衡与公平性违规之间的形式等价关系。理论预测,在特征重叠假设下,强度为α的虚假相关,等效于子群体不平衡比r≈(1+α)/(1-α)导致的最差组准确率下降。在六个数据集和三种架构上的实证验证表明,预测的等价性在最差组准确率3%误差范围内成立,支持了跨问题领域对治偏方法的合理转移。本工作将公平性、鲁棒性与分布偏移文献统一于同一视角。
原文摘要 · Abstract (English)
Machine learning systems exhibit diverse failure modes: unfairness toward protected groups, brittleness to spurious correlations, poor performance on minority sub-populations, which are typically studied in isolation by distinct research communities. We propose a unifying theoretical framework that characterizes when different bias mechanisms produce quantitatively equivalent effects on model performance. By formalizing biases as violations of conditional independence through information-theoretic measures, we prove formal equivalence conditions relating spurious correlations, subpopulation shift, class imbalance, and fairness violations. Our theory predicts that a spurious correlation of strength $α$ produces equivalent worst-group accuracy degradation as a sub-population imbalance ratio $r \approx (1+α)/(1-α)$ under feature overlap assumptions. Empirical validation in six datasets and three architectures confirms that predicted equivalences hold within the accuracy of the worst group 3\%, enabling the principled transfer of debiasing methods across problem domains. This work bridges the literature on fairness, robustness, and distribution shifts under a common perspective.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。