arXiv:2504.17314cs.LGcs.CV2025-04被引 1

通过平衡类别条件分布,无需标注或预测偏见即可消除模型错误关联。

Class-Conditional Distribution Balancing for Group Robust Classification

  • 将虚假相关归因于类别条件分布失衡,设计重加权策略实现自动平衡。
  • 在多个数据集上达到顶尖性能,媲美依赖偏见标注的方法。
  • 适合资源有限的罕见领域,无需额外标注或大模型辅助。

虚假相关导致模型因错误原因正确预测,严重威胁实际应用中的鲁棒泛化能力。现有研究将其归因于群体不平衡,通过最大化群体平衡准确率或最差群体准确率来解决,但高度依赖昂贵的偏见标注。一种折中方案是利用大规模预训练基础模型预测偏见信息,却需海量数据,在资源受限的稀有领域不切实际。为此,本文提出新视角:将虚假相关视为类别条件分布的不平衡或错配,提出一种简单而有效的鲁棒学习方法,完全无需偏见标注或预测。目标是最大化标签在虚假因素给定下的条件熵(不确定性),通过样本重加权实现类别条件分布平衡,自动突出少数群体与类别,有效消除虚假相关,生成无偏的数据分布用于分类。大量实验与分析表明,该方法持续表现优异,性能媲美依赖偏见监督的方法。

原文摘要 · Abstract (English)

Spurious correlations that lead models to correct predictions for the wrong reasons pose a critical challenge for robust real-world generalization. Existing research attributes this issue to group imbalance and addresses it by maximizing group-balanced or worst-group accuracy, which heavily relies on expensive bias annotations. A compromise approach involves predicting bias information using extensively pretrained foundation models, which requires large-scale data and becomes impractical for resource-limited rare domains. To address these challenges, we offer a novel perspective by reframing the spurious correlations as imbalances or mismatches in class-conditional distributions, and propose a simple yet effective robust learning method that eliminates the need for both bias annotations and predictions. With the goal of maximizing the conditional entropy (uncertainty) of the label given spurious factors, our method leverages a sample reweighting strategy to achieve class-conditional distribution balancing, which automatically highlights minority groups and classes, effectively dismantling spurious correlations and producing a debiased data distribution for classification. Extensive experiments and analysis demonstrate that our approach consistently delivers state-of-the-art performance, rivaling methods that rely on bias supervision.

鲁棒学习虚假相关分布平衡无监督

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。