arXiv:2501.10453cs.LGcs.AI2025-01被引 5

提出新方法检测并减轻大模型中的隐性偏见,尤其关注多重社会属性交叉时的歧视问题。

Uncovering Bias in Foundation Models: Impact, Testing, Harm, and Mitigation

  • 设计语义探针系统测试大模型在性别、种族等属性上的显性和隐性偏见。
  • 发现模型在性别×种族等交叉属性上存在复杂且普遍的偏见,影响公平性。
  • 提出自适应概率调整技术,无需重训练即可显著提升模型公平性,适合从业者部署。

基础模型(FMs)在涵盖社会与历史知识的海量数据上训练,其固有的偏见对医疗、教育、金融等领域公平性构成重大挑战。这些偏见源于训练数据中对刻板印象和不平等现象的过度反映,加剧了现实歧视、强化有害刻板印象并削弱公众对AI的信任。为此,我们提出三重探针测试(TriProTesting),一种基于语义设计探针的系统化检测方法,可识别显性和隐性偏见。实验表明,包括CLIP、ALIGN、BridgeTower和OWLv2在内的多个基础模型在单一及混合社会属性(性别、种族、年龄、职业)上普遍存在偏见。特别地,在性别×种族、性别×年龄、性别×职业等交叉组合中发现了复杂的混合偏见,揭示出更深层的歧视机制。我们进一步提出自适应逻辑调整(AdaLogAdjustment),一种后处理技术,通过动态重新分配概率权重有效缓解偏见,显著提升公平性而无需重新训练模型。研究强调亟需伦理AI实践与跨学科解决方案,从模型层面延伸至社会结构。本工作提供了一种可扩展、可解释的公平性增强方案,为未来公平AI技术研究提供实用洞见。

原文摘要 · Abstract (English)

Bias in Foundation Models (FMs) - trained on vast datasets spanning societal and historical knowledge - poses significant challenges for fairness and equity across fields such as healthcare, education, and finance. These biases, rooted in the overrepresentation of stereotypes and societal inequalities in training data, exacerbate real-world discrimination, reinforce harmful stereotypes, and erode trust in AI systems. To address this, we introduce Trident Probe Testing (TriProTesting), a systematic testing method that detects explicit and implicit biases using semantically designed probes. Here we show that FMs, including CLIP, ALIGN, BridgeTower, and OWLv2, demonstrate pervasive biases across single and mixed social attributes (gender, race, age, and occupation). Notably, we uncover mixed biases when social attributes are combined, such as gender x race, gender x age, and gender x occupation, revealing deeper layers of discrimination. We further propose Adaptive Logit Adjustment (AdaLogAdjustment), a post-processing technique that dynamically redistributes probability power to mitigate these biases effectively, achieving significant improvements in fairness without retraining models. These findings highlight the urgent need for ethical AI practices and interdisciplinary solutions to address biases not only at the model level but also in societal structures. Our work provides a scalable and interpretable solution that advances fairness in AI systems while offering practical insights for future research on fair AI technologies.

偏见检测公平性大模型后处理

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。