arXiv:2411.00360cs.LGcs.CV2024-11NeurIPS被引 5

通过误标样本视角,简单有效缓解数据偏见问题。

A Simple Remedy for Dataset Bias via Self-Influence: A Mislabeled Sample Perspective

  • 利用影响函数识别误标样本,间接发现偏见冲突样本。
  • 在多个数据集上提升检测精度,显著改善模型公平性。
  • 无需先验知识,可与现有去偏方法协同增效。

从有偏数据中学习泛化模型是实现深度学习公平性的关键任务。现有研究试图在无偏知识或无偏数据集的情况下,识别并利用不受虚假相关性干扰的偏见冲突样本。然而,虚假相关性仍是主要挑战,根源在于难以精准检测这些样本。本文受误标样本与偏见冲突样本相似性的启发,从误标样本检测的新视角切入,深入分析影响函数这一标准误标检测方法,提出一种简单而有效的偏见修正策略。在多个数据集上的综合分析与实验表明,该新视角能显著提升检测精度,并有效修正偏见模型。此外,该方法与现有技术具有互补性,即使应用于已进行近期去偏处理的模型,仍能带来性能提升。

原文摘要 · Abstract (English)

Learning generalized models from biased data is an important undertaking toward fairness in deep learning. To address this issue, recent studies attempt to identify and leverage bias-conflicting samples free from spurious correlations without prior knowledge of bias or an unbiased set. However, spurious correlation remains an ongoing challenge, primarily due to the difficulty in precisely detecting these samples. In this paper, inspired by the similarities between mislabeled samples and bias-conflicting samples, we approach this challenge from a novel perspective of mislabeled sample detection. Specifically, we delve into Influence Function, one of the standard methods for mislabeled sample detection, for identifying bias-conflicting samples and propose a simple yet effective remedy for biased models by leveraging them. Through comprehensive analysis and experiments on diverse datasets, we demonstrate that our new perspective can boost the precision of detection and rectify biased models effectively. Furthermore, our approach is complementary to existing methods, showing performance improvement even when applied to models that have already undergone recent debiasing techniques.

去偏影响函数公平性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。