不依赖标签,自动发现隐藏群体并提升最差群体的公平性
Robust Mixture Models for Algorithmic Fairness Under Latent Heterogeneity
- 通过混合模型自动识别数据中隐含的群体结构
- 在真实数据集上显著改善最差群体的表现,平均性能仍保持竞争力
- 适合未知或动态变化的不公平来源场景
以平均性能优化的标准机器学习模型在少数群体上表现往往不佳,且对分布偏移缺乏鲁棒性。当群体特征为隐性,且连续与离散特征存在复杂交互时,该问题更为严重。本文提出ROME(RObust Mixture Ensemble)框架,从数据中学习隐含群体结构的同时优化最差群体表现。ROME采用两种方法:针对线性模型使用期望最大化算法,针对非线性场景采用神经网络混合专家模型。通过模拟和真实数据集实验,验证了ROME在算法公平性方面显著优于传统方法,同时保持有竞争力的平均性能。关键优势在于无需预先定义群体标签,适用于不公平来源未知或动态变化的实际场景。
原文摘要 · Abstract (English)
Standard machine learning models optimized for average performance often fail on minority subgroups and lack robustness to distribution shifts. This challenge worsens when subgroups are latent and affected by complex interactions among continuous and discrete features. We introduce ROME (RObust Mixture Ensemble), a framework that learns latent group structure from data while optimizing for worst-group performance. ROME employs two approaches: an Expectation-Maximization algorithm for linear models and a neural Mixture-of-Experts for nonlinear settings. Through simulations and experiments on real-world datasets, we demonstrate that ROME significantly improves algorithmic fairness compared to standard methods while maintaining competitive average performance. Importantly, our method requires no predefined group labels, making it practical when sources of disparities are unknown or evolving.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。