arXiv:2411.18442cs.LGcs.AI2024-11被引 1

用多样性引导的自训练缓解数据偏差,提升模型公平性。

Metric-DST: Mitigating Selection Bias Through Diversity-Guided Semi-Supervised Metric Learning

  • 通过度量学习空间选择多样样本,避免仅选高置信度数据
  • 在合成与真实数据集上均显著降低偏差,提升模型鲁棒性
  • 适用于有偏差的数据场景,尤其适合追求公平性的研究者

选择偏差是机器学习公平性中的关键挑战,因训练数据对人群代表性不足,可能导致对少数群体表现不佳。半监督学习如自训练可通过引入无标签数据来理解整体分布,缓解偏差。然而传统自训练仅选择高置信度样本,可能强化已有偏差。本文提出Metric-DST,一种基于多样性的自训练策略,利用度量学习的隐式嵌入空间,主动纳入更多样化的样本以对抗信心驱动的偏差。该方法在具有人为引入偏差的生成数据和真实世界数据集,以及存在固有偏差的分子生物学预测任务中,均学习到更稳健的模型。Metric-DST提供了一种灵活、普适的解决方案,可有效缓解选择偏差,提升模型公平性。

原文摘要 · Abstract (English)

Selection bias poses a critical challenge for fairness in machine learning, as models trained on data that is less representative of the population might exhibit undesirable behavior for underrepresented profiles. Semi-supervised learning strategies like self-training can mitigate selection bias by incorporating unlabeled data into model training to gain further insight into the distribution of the population. However, conventional self-training seeks to include high-confidence data samples, which may reinforce existing model bias and compromise effectiveness. We propose Metric-DST, a diversity-guided self-training strategy that leverages metric learning and its implicit embedding space to counter confidence-based bias through the inclusion of more diverse samples. Metric-DST learned more robust models in the presence of selection bias for generated and real-world datasets with induced bias, as well as a molecular biology prediction task with intrinsic bias. The Metric-DST learning strategy offers a flexible and widely applicable solution to mitigate selection bias and enhance fairness of machine learning models.

公平性自训练度量学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。