DCAST通过多样本自训练缓解分类偏见,提升模型公平性。
DCAST: Diverse Class-Aware Self-Training Mitigates Selection Bias for Fairer Learning
- 引入多样化类感知自训练,增强未标注数据利用
- 在11个数据集上显著降低层级偏见,多分类效果更优
- 适合需要公平性保障的视觉与生物医学场景
机器学习中的公平性旨在减少因性别、年龄等敏感特征导致的模型偏见,此类偏见常源于训练数据中人群表示不均的选取偏差。值得注意的是,与敏感特征无关的偏见难以识别,尤其在计算机视觉和分子生物医学等高维复杂数据中尤为突出。现有方法对未识别偏见的缓解与评估仍严重不足。本文提出:(i) 多样化类感知自训练(DCAST),一种无需模型依赖的偏见缓解策略,能感知类别特异性偏见,在提升样本多样性以对抗传统自训练的确认偏见的同时,利用未标注样本优化底层群体表征;(ii) 层级偏见,一种无需先验知识的多变量、类感知偏见生成机制。使用DCAST训练的模型在11个数据集上展现出对层级及其他偏见更强的鲁棒性,优于传统自训练和六种主流领域自适应技术,多分类任务中优势最为显著,凸显其在多种场景下实现更公平学习的潜力。
原文摘要 · Abstract (English)
Fairness in machine learning seeks to mitigate model bias against individuals based on sensitive features such as sex or age, often caused by an uneven representation of the population in the training data due to selection bias. Notably, bias unascribed to sensitive features is challenging to identify and typically goes undiagnosed, despite its prominence in complex high-dimensional data from fields like computer vision and molecular biomedicine. Strategies to mitigate unidentified bias and evaluate mitigation methods are crucially needed, yet remain underexplored. We introduce: (i) Diverse Class-Aware Self-Training (DCAST), model-agnostic mitigation aware of class-specific bias, which promotes sample diversity to counter confirmation bias of conventional self-training while leveraging unlabeled samples for an improved representation of the underlying population; (ii) hierarchy bias, multivariate and class-aware bias induction without prior knowledge. Models learned with DCAST showed improved robustness to hierarchy and other biases across eleven datasets, against conventional self-training and six prominent domain adaptation techniques. Advantage was largest on multi-class classification, emphasizing DCAST as a promising strategy for fairer learning in different contexts.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。