用强化学习与思维链提升大模型识别并减轻性别偏见的能力。
Detection, Classification, and Mitigation of Gender Bias in Large Language Models
- 通过思维链分步推理,增强对复杂偏见问题的理解能力。
- 在三个子任务中均排名第一,显著提升偏见检测与分类准确率。
- 适用于关注大模型公平性、伦理安全的研究者与开发者。
随着大语言模型(LLMs)的快速发展,其在多个领域显著提升了效率。然而,近期研究发现,LLMs常表现出性别偏见,带来严重的社会影响。因此,检测、分类和缓解性别偏见成为关键研究方向。在NLPCC 2025共享任务7:中文性别偏见检测、分类与缓解挑战中,我们探索如何提升LLMs在性别偏见检测、分类与缓解方面的能力。针对子任务1和2,采用链式思维(CoT)推理,通过分阶段多步思考简化复杂偏见查询,提升响应准确性;针对子任务3,基于强化学习构建偏好数据集,使用GPT-4标注,并应用直接偏好优化(DPO)引入损失函数,明确偏好无偏完成结果。该方法在所有三个子任务中均取得第一名。
原文摘要 · Abstract (English)
With the rapid development of large language models (LLMs), they have significantly improved efficiency across a wide range of domains. However, recent studies have revealed that LLMs often exhibit gender bias, leading to serious social implications. Detecting, classifying, and mitigating gender bias in LLMs has therefore become a critical research focus. In the NLPCC 2025 Shared Task 7: Chinese Corpus for Gender Bias Detection, Classification and Mitigation Challenge, we investigate how to enhance the capabilities of LLMs in gender bias detection, classification, and mitigation. We adopt reinforcement learning, chain-of-thoughts (CoT) reasoning, and supervised fine-tuning to handle different Subtasks. Specifically, for Subtasks 1 and 2, we leverage the internal reasoning capabilities of LLMs to guide multi-step thinking in a staged manner, which simplifies complex biased queries and improves response accuracy. For Subtask 3, we employ a reinforcement learning-based approach, annotating a preference dataset using GPT-4. We then apply Direct Preference Optimization (DPO) to mitigate gender bias by introducing a loss function that explicitly favors less biased completions over biased ones. Our approach ranked first across all three subtasks of the NLPCC 2025 Shared Task 7.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。