arXiv:2412.10575cs.LGcs.AI2024-12AAAI被引 2

用多校准重新评估数据增强,发现公平混元反而降低公平性

Who's the (Multi-)Fairest of Them All: Rethinking Interpolation-Based Data Augmentation Through the Lens of Multicalibration

  • 以多校准为标准,测试四种公平混元方法的公平性表现
  • 公平混元在81个边缘群体中普遍恶化公平性和准确率
  • 普通混元+事后多校准可更好提升小群体公平性

数据增强方法,尤其是当前最先进的基于插值的方法如公平混元(Fair Mixup),已被广泛证明能提升模型公平性。然而,这些公平性评估仅基于不反映模型不确定性的指标,且仅针对单一、相对较大的少数群体数据集。为此,多校准(multicalibration)被引入,用于在考虑不确定性的同时衡量公平性,并处理多个少数群体。但现有改进多校准的方法需减少初始训练数据以创建保留集进行事后处理,这在少数群体训练数据本就稀疏时并不理想。本文利用多校准更严格地检验分类任务中的数据增强公平性。我们在两个结构化数据分类问题上对四种公平混元版本进行压力测试,涉及最多81个被边缘化的群体,评估多校准违反情况与平衡准确率。结果发现,在几乎所有实验中,公平混元均恶化了基线性能和公平性;而简单的原始混元(vanilla Mixup)反而优于公平混元和基线,尤其在小群体校准中表现更佳。将原始混元与多校准事后处理结合(通过保留集强制执行多校准),进一步提升了公平性。

原文摘要 · Abstract (English)

Data augmentation methods, especially SoTA interpolation-based methods such as Fair Mixup, have been widely shown to increase model fairness. However, this fairness is evaluated on metrics that do not capture model uncertainty and on datasets with only one, relatively large, minority group. As a remedy, multicalibration has been introduced to measure fairness while accommodating uncertainty and accounting for multiple minority groups. However, existing methods of improving multicalibration involve reducing initial training data to create a holdout set for post-processing, which is not ideal when minority training data is already sparse. This paper uses multicalibration to more rigorously examine data augmentation for classification fairness. We stress-test four versions of Fair Mixup on two structured data classification problems with up to 81 marginalized groups, evaluating multicalibration violations and balanced accuracy. We find that on nearly every experiment, Fair Mixup \textit{worsens} baseline performance and fairness, but the simple vanilla Mixup \textit{outperforms} both Fair Mixup and the baseline, especially when calibrating on small groups. \textit{Combining} vanilla Mixup with multicalibration post-processing, which enforces multicalibration through post-processing on a holdout set, further increases fairness.

公平性数据增强多校准混元

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。