arXiv:2511.06023cs.CL2025-11被引 2

用多维度奖励优化,让大模型摆脱中文语境下的歧视偏见。

Multi-Reward GRPO Fine-Tuning for De-biasing Large Language Models: A Study Based on Chinese-Context Discrimination Data

  • 构建中文歧视类别数据集,用双响应对比训练多维奖励模型。
  • 多奖励信号下,模型偏见强度显著下降,流畅性不受影响。
  • 适合关注文化敏感性对齐与伦理调优的研究者参考。

大型语言模型常表现出隐含偏见与歧视倾向,反映社会刻板印象。尽管近期的对齐技术如RLHF和DPO已缓解部分问题,但在应对文化特定、多维度歧视方面仍有限。本文提出一种多奖励组相对策略优化(Multi-Reward GRPO)框架,用于微调大模型以实现伦理化与无偏行为。方法基于中文语境下的区域、民族与职业歧视类别,构建合成英文数据集,每条样本配以中立与有偏响应,训练基于DeBERTa-v3的奖励模型,提供公平性、中立性与语言质量的多维奖励信号。该模型指导GRPO微调,优化输出在伦理维度的表现。实验显示,偏见强度显著降低,且未损害流畅性与信息量。本研究验证了基于多奖励的GRPO在去偏方面的有效性,为文化情境下的伦理对齐提供了可复现框架。

原文摘要 · Abstract (English)

Large Language Models (LLMs) often exhibit implicit biases and discriminatory tendencies that reflect underlying social stereotypes. While recent alignment techniques such as RLHF and DPO have mitigated some of these issues, they remain limited in addressing culturally specific and multi-dimensional forms of discrimination. This paper proposes a Multi-Reward Group Relative Policy Optimization (GRPO) framework to fine-tune LLMs toward ethical and bias-free behavior. Our approach constructs a synthetic English-language dataset derived from Chinese-context discrimination categories, including regional, ethnic, and occupational biases. Each instance is paired with both neutral and biased responses to train a reward model based on DeBERTa-v3, which provides multi-dimensional reward signals capturing fairness, neutrality, and linguistic quality. The trained reward model then guides GRPO fine-tuning to optimize model outputs along these ethical dimensions. Experimental results demonstrate significant reductions in bias intensity and improved alignment with non-discriminatory standards without compromising fluency or informativeness. This study highlights the effectiveness of GRPO-based multi-reward optimization for de-biasing LLMs and offers a replicable framework for cultural-contextual ethical alignment.

大模型去偏伦理对齐多奖励学习中文语境

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。