arXiv:2509.08942cs.LG2025-09被引 1

考虑组内分布不确定性,提升模型在少数群体上的鲁棒性。

Group Distributionally Robust Machine Learning under Group Level Distributional Uncertainty

  • 用Wasserstein距离建模每组内部的分布不确定性和最差组优化
  • 在真实数据集上显著改善少数群体的性能表现
  • 适合处理数据噪声大、分布动态变化的场景

机器学习模型性能高度依赖训练数据的质量与代表性。在存在多个异构数据生成源的应用中,标准方法常学习到虚假关联,虽平均表现良好,却在异常或未充分代表的群体上性能下降。已有工作通过优化最差组表现来缓解此问题,但通常假设每组的底层分布可由训练数据准确估计,这一条件在噪声大、非平稳和动态变化的环境中往往不成立。本文提出一种新框架,基于Wasserstein分布鲁棒优化(DRO),同时考虑组内分布不确定性并保持提升最差组性能的目标。我们设计了一种梯度下降-上升算法求解该DRO问题,并提供收敛性分析。最后在真实数据集上验证了方法的有效性。

原文摘要 · Abstract (English)

The performance of machine learning (ML) models critically depends on the quality and representativeness of the training data. In applications with multiple heterogeneous data generating sources, standard ML methods often learn spurious correlations that perform well on average but degrade performance for atypical or underrepresented groups. Prior work addresses this issue by optimizing the worst-group performance. However, these approaches typically assume that the underlying data distributions for each group can be accurately estimated using the training data, a condition that is frequently violated in noisy, non-stationary, and evolving environments. In this work, we propose a novel framework that relies on Wasserstein-based distributionally robust optimization (DRO) to account for the distributional uncertainty within each group, while simultaneously preserving the objective of improving the worst-group performance. We develop a gradient descent-ascent algorithm to solve the proposed DRO problem and provide convergence results. Finally, we validate the effectiveness of our method on real-world data.

分布鲁棒公平学习最差组优化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。