arXiv:2503.17632cs.LGcs.AI2025-03

让模型学会对有偏数据保持不确定,提升泛化能力

FairFlow: Mitigating Dataset Biases through Undecided Learning

  • 通过扰动数据和模型生成多种有偏视图
  • 在域外和困难样本上性能显著提升
  • 适合需要抗偏见的高可靠性场景

语言模型易受数据偏差影响,产生捷径依赖和虚假相关,导致在新数据上表现下降。我们提出新的去偏框架 FairFlow,通过学习对已知或未知偏差的数据样本或表示保持不确定来缓解偏差。该框架引入两类关键组件:一组数据和模型扰动操作,用于生成输入样本的不同有偏视图;以及一种对比目标函数,从这些有偏视图中学习去偏且鲁棒的表示。实验表明,FairFlow 在多项基准测试中优于现有去偏方法,尤其在域外数据和难例测试集上表现更优,同时不牺牲域内性能。

原文摘要 · Abstract (English)

Language models are prone to dataset biases, known as shortcuts and spurious correlations in data, which often result in performance drop on new data. We present a new debiasing framework called ``FairFlow'' that mitigates dataset biases by learning to be undecided in its predictions for data samples or representations associated with known or unknown biases. The framework introduces two key components: a suite of data and model perturbation operations that generate different biased views of input samples, and a contrastive objective that learns debiased and robust representations from the resulting biased views of samples. Experiments show that FairFlow outperforms existing debiasing methods, particularly against out-of-domain and hard test samples without compromising the in-domain performance

去偏语言模型鲁棒性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。