arXiv:2512.09854cs.CL2025-12

用偏好模型筛选和迭代修正,减少英乌语大模型的社交偏见。

Mitigating Social Bias in English and Urdu Language Models Using PRM-Guided Candidate Selection and Sequential Refinement

  • 基于偏好模型筛选候选输出并逐步优化,不需重新训练。
  • 英乌双语测试中偏见显著降低,但乌尔都语公平性仍较差。
  • 方法可解释、可扩展,适合低资源语言公平性研究。

大型语言模型(LLMs)在人类交流、决策支持、内容生成和信息检索中日益扮演关键角色。尽管表现出色,这些系统在面对社会敏感语言时常产生偏见或刻板内容,尤其在低资源语言中更为严重,因训练数据有限且文化代表性不足。本文开展了一项推理阶段偏见缓解的综合性研究,该策略无需重训练或微调,直接作用于模型输出。基于偏好排名模型(PRMs),我们构建统一评估框架,对比三种方法:(1) 基线单词生成,(2) 基于PRM的best-of-N采样选择,(3) 基于PRM批判的序列化迭代优化。我们在200个英文提示及其乌尔都语对应版本上评估,涵盖性别、种族、宗教、国籍、残疾、职业、年龄和经济状况等社会文化情境。使用GPT-3.5作为候选生成器,GPT-4o-mini作为PRM驱动的偏见与效用评分器,我们对偏见降低、效用保持及跨语言差异进行了量化分析。结果表明:(a) 两种语言均显著优于基线;(b) 所有方法下乌尔都语公平性得分始终更低,揭示多语言模型训练中的结构性不平等;(c) PRM-Select与PRM-Sequential表现出不同的改进轨迹。本研究贡献了可扩展的方法、可解释的度量标准及跨语言比较,为未来低资源语言公平性评估提供支持。

原文摘要 · Abstract (English)

Large language models (LLMs) increasingly mediate human communication, decision support, content creation, and information retrieval. Despite impressive fluency, these systems frequently produce biased or stereotypical content, especially when prompted with socially sensitive language. A growing body of research has demonstrated that such biases disproportionately affect low-resource languages, where training data is limited and culturally unrepresentative. This paper presents a comprehensive study of inference-time bias mitigation, a strategy that avoids retraining or fine-tuning and instead operates directly on model outputs. Building on preference-ranking models (PRMs), we introduce a unified evaluation framework comparing three methods: (1) baseline single-word generation, (2) PRM-Select best-of-N sampling, and (3) PRM-Sequential refinement guided by PRM critiques. We evaluate these techniques across 200 English prompts and their Urdu counterparts, designed to reflect socio-cultural contexts relevant to gender, ethnicity, religion, nationality, disability, profession, age, and socioeconomic categories. Using GPT-3.5 as a candidate generator and GPT-4o-mini as a PRM-based bias and utility scorer, we provide an extensive quantitative analysis of bias reduction, utility preservation, and cross-lingual disparities. Our findings show: (a) substantial gains over the baseline for both languages; (b) consistently lower fairness scores for Urdu across all methods, highlighting structural inequities in multilingual LLM training; and (c) distinct improvement trajectories between PRM-Select and PRM-Sequential. The study contributes an extensible methodology, interpretable metrics, and cross-lingual comparisons that can support future work on fairness evaluation in low-resource languages.

偏见缓解多语言推理优化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。