用多目标优化解决多个数据偏见共存时的模型偏差问题
Improving Robustness to Multiple Spurious Correlations by Multi-Objective Optimization
- 按数据分组识别不同偏见,动态调整损失权重以平衡冲突
- 在三个多偏见数据集上性能领先,单偏见数据集也表现优异
- 提出全新多偏见基准数据集MultiCelebA,推动公平性研究
我们研究在存在多个数据偏见的场景下训练无偏且准确模型的问题。该问题极具挑战性,因为多个偏见会导致训练中出现多重不良捷径,且缓解一个偏见可能加剧另一个。为此,我们提出一种新颖的训练方法:首先将训练数据分组,使每组引发不同捷径;然后优化各组损失的线性组合,并动态调整权重以缓解组间性能冲突;该方法基于多目标优化理论,旨在实现最小最大帕累托解。我们还构建了一个包含多个偏见的新基准数据集MultiCelebA,用于在真实且复杂的场景下评估去偏训练方法。实验表明,我们的方法在三个多偏见数据集上均取得最佳表现,并在传统单偏见数据集上也展现出优越性能。
原文摘要 · Abstract (English)
We study the problem of training an unbiased and accurate model given a dataset with multiple biases. This problem is challenging since the multiple biases cause multiple undesirable shortcuts during training, and even worse, mitigating one may exacerbate the other. We propose a novel training method to tackle this challenge. Our method first groups training data so that different groups induce different shortcuts, and then optimizes a linear combination of group-wise losses while adjusting their weights dynamically to alleviate conflicts between the groups in performance; this approach, rooted in the multi-objective optimization theory, encourages to achieve the minimax Pareto solution. We also present a new benchmark with multiple biases, dubbed MultiCelebA, for evaluating debiased training methods under realistic and challenging scenarios. Our method achieved the best on three datasets with multiple biases, and also showed superior performance on conventional single-bias datasets.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。