arXiv:2510.06354cs.CL2025-10EMNLP被引 8

用目标分布定义大模型偏见,实现更真实或更平等的输出对齐。

LLM Bias Detection and Mitigation through the Lens of Desired Distributions

  • 基于目标分布设计加权自适应损失,微调大模型输出
  • 在真实职业分布下减少30%-75%偏见,平等设定下接近完全消除
  • 适合关注模型输出真实性或社会公平性的研究者与应用方

以往偏见缓解工作多聚焦于促进社会平等和人口均等性,较少关注使大模型输出与目标分布对齐。例如,我们可能希望模型输出符合现实世界的职业分布以增强事实准确性。因此,本文将偏见定义为输出分布偏离目标分布的程度,目标分布可为均等分布或真实分布,具体取决于应用场景。提出一种基于加权自适应损失的微调方法,使大模型在性别-职业输出分布上与目标分布对齐,同时保持语言建模能力。利用美国劳工统计局2024年数据构建三类职业集(男性主导、女性主导、性别平衡),评估该方法在反映现实与追求平等两种情形下的表现。在三种掩码语言模型上均观测到偏见。在平等目标下实现近完全缓解,在真实分布目标下实现30%-75%的偏见降低。自回归大模型在平等目标下无偏见,但在真实分布下仍存在显著偏见,其中Llama Instruct模型(3.2-3B、3.1-8B)实现50%-62%的偏见减少。

原文摘要 · Abstract (English)

Although prior work on bias mitigation has focused on promoting social equality and demographic parity, less attention has been given to aligning LLM's outputs to desired distributions. For example, we might want to align a model with real-world distributions to support factual grounding. Thus, we define bias as deviation from a desired distribution, which may be an equal or real-world distribution, depending on application goals. We propose a weighted adaptive loss based fine-tuning method that aligns LLM's gender-profession output distribution with the desired distribution, while preserving language modeling capability. Using 3 profession sets -- male-dominated, female-dominated, and gender-balanced -- derived from U.S. labor statistics (2024), we assess both our adaptive method for reflecting reality and a non-adaptive variant for equality. Across three masked language models, bias is observed under both distributions. We achieve near-complete mitigation under equality and 30-75% reduction under real-world settings. Autoregressive LLMs show no bias under equality but notable bias under real-world settings, with the Llama Instruct models (3.2-3B, 3.1-8B) achieving a 50-62% reduction.

偏见检测分布对齐语言模型公平性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。