研究数据筛选如何影响模型偏差,发现选难样本未必提升鲁棒性
The Impact of Coreset Selection on Spurious Correlations and Group Robustness
- 用嵌入特征评分选样本比基于学习动态的选法更少加剧偏差
- 优先选难样本可降低偏差水平但无法保证下游模型鲁棒性
- 首次系统分析数据缩减对虚假相关性和群体鲁棒性的影响
数据集缩减方法在保持模型性能的同时减少训练数据量,但在存在偏见的数据集中,模型可能学习到虚假相关而非因果特征。本文首次系统分析数据选择对所选核心集虚假偏差水平及下游模型鲁棒性的影响。实验覆盖10个虚假相关性基准、5种样本重要性/难度评分指标和5种数据选择策略,涵盖广泛的核心集规模。结果揭示了样本难度与偏差一致性之间复杂的相互作用,以及数据偏差与模型鲁棒性的关联。例如,基于嵌入特征的样本评分方法比基于学习动态的方法更不易无意中加剧偏差;更重要的是,尽管某些方法通过优先选择难样本能降低偏差,却无法可靠保证下游模型的鲁棒性。
原文摘要 · Abstract (English)
Coreset selection methods have shown promise in reducing the training data size while maintaining model performance for data-efficient machine learning. However, as many datasets suffer from biases that cause models to learn spurious correlations instead of causal features, it is important to understand whether and how dataset reduction methods may perpetuate, amplify, or mitigate these biases. In this work, we conduct the first comprehensive analysis of the implications of data selection on the spurious bias levels of the selected coresets and the robustness of downstream models trained on them. We use an extensive experimental setting spanning ten different spurious correlations benchmarks, five score metrics to characterize sample importance/ difficulty, and five data selection policies across a broad range of coreset sizes. Thereby, we unravel a series of nontrivial nuances in interactions between sample difficulty and bias alignment, as well as dataset bias and resultant model robustness. For example, we find that selecting coresets using embedding-based sample characterization scores runs a comparatively lower risk of inadvertently exacerbating bias than selecting using characterizations based on learning dynamics. Most importantly, our analysis reveals that although some coreset selection methods could achieve lower bias levels by prioritizing difficult samples, they do not reliably guarantee downstream robustness.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。