arXiv:2505.06831cs.CV2025-05被引 1

通过细粒度分布平衡,提升无偏标注场景下的模型鲁棒性。

Fine-Grained Class-Conditional Distribution Balancing for Debiased Learning

  • 用混淆矩阵细化类别条件分布描述,替代单一高斯近似。
  • 在多类和多重捷径场景中,性能超越有监督偏见方法。
  • 无需偏见标注,适合真实世界中缺乏标注的公平学习任务。

在缺乏偏见标注的情况下,实现群体鲁棒泛化仍是重大挑战。现有类别条件分布平衡(CCDB)方法通过无偏见方式匹配分布,取得良好效果,但其使用单一高斯近似每种分布,过于简略,难以适用于真实场景。为此,本文提出多阶段数据选择性重训练策略(MST),利用硬混淆矩阵更细致地刻画分布。在此基础上,提出细粒度CCDB(FG-CCDB),通过混淆单元级别的重加权实现更精确的分布匹配。FG-CCDB从全局学习样本权重,有效缓解虚假相关,且存储与计算开销低。大量实验表明,MST可作为真实偏见标注的可靠代理,能无缝集成至有监督方法中;结合FG-CCDB后,在二分类任务上表现媲美有监督方法,在高度偏见的多分类和多捷径场景中显著优于后者。

原文摘要 · Abstract (English)

Achieving group-robust generalization in the presence of spurious correlations remains a significant challenge, particularly when bias annotations are unavailable. Recent studies on Class-Conditional Distribution Balancing (CCDB) reveal that spurious correlations often stem from mismatches between the class-conditional and marginal distributions of bias attributes. They achieve promising results by addressing this issue through simple distribution matching in a bias-agnostic manner. However, CCDB approximates each distribution using a single Gaussian, which is overly simplistic and rarely holds in real-world applications. To address this limitation, we propose a novel Multi-stage data-Selective reTraining strategy (MST), which describes each distribution in greater detail using the hard confusion matrix. Building on these finer descriptions, we propose a fine-grained variant of CCDB, termed FG-CCDB, which enhances distribution matching through more precise confusion-cell-wise reweighting. FG-CCDB learns sample weights from a global perspective, effectively mitigating spurious correlations without incurring substantial storage or computational overhead. Extensive experiments demonstrate that MST serves as a reliable proxy for ground-truth bias annotations and can be seamlessly integrated with bias-supervised methods. Moreover, when combined with FG-CCDB, our method performs on par with bias-supervised approaches on binary classification tasks and significantly outperforms them in highly biased multi-class and multi-shortcut scenarios.

去偏学习分布平衡细粒度建模

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。