arXiv:2503.07315cs.LGcs.AI2025-03ICLR被引 8

用影响函数重加权无标签数据,提升模型对子群体分布偏移的鲁棒性

Group-robust Sample Reweighting for Subpopulation Shifts via Influence Functions

  • 先学无标签数据表示,再用影响函数重加权并微调最后层
  • 仅用少量有标签子群体数据,性能超越需更多标签的先进方法
  • 适合标注成本高、子群体分布易变的实际场景

机器学习模型在数据分布的子群体间表现不均,当部署时子群体比例发生变化时,模型泛化能力面临挑战。现有方法通常依赖大量带子群体标签的数据进行训练或超参数调优,以最小化最差子群体的损失,但高质量标签成本高昂。为此,本文提出一种新范式:利用有限的子群体标签数据作为目标,优化其他无标签数据的权重。我们提出群组鲁棒样本重加权(GSR),采用两阶段策略:首先从无标签数据中学习表示,然后通过影响函数迭代重加权并重新训练模型最后一层。GSR理论严谨、实现轻量,能有效提升对子群体分布偏移的鲁棒性。实验表明,其性能优于同等或更高标签用量的现有最优方法。

原文摘要 · Abstract (English)

Machine learning models often have uneven performance among subpopulations (a.k.a., groups) in the data distributions. This poses a significant challenge for the models to generalize when the proportions of the groups shift during deployment. To improve robustness to such shifts, existing approaches have developed strategies that train models or perform hyperparameter tuning using the group-labeled data to minimize the worst-case loss over groups. However, a non-trivial amount of high-quality labels is often required to obtain noticeable improvements. Given the costliness of the labels, we propose to adopt a different paradigm to enhance group label efficiency: utilizing the group-labeled data as a target set to optimize the weights of other group-unlabeled data. We introduce Group-robust Sample Reweighting (GSR), a two-stage approach that first learns the representations from group-unlabeled data, and then tinkers the model by iteratively retraining its last layer on the reweighted data using influence functions. Our GSR is theoretically sound, practically lightweight, and effective in improving the robustness to subpopulation shifts. In particular, GSR outperforms the previous state-of-the-art approaches that require the same amount or even more group labels.

子群体偏移影响函数样本重加权小样本学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。