arXiv:2409.04340cs.LGcs.AI2024-09被引 17

为大模型年龄偏见设计公平奖励机制,提升跨年龄响应一致性

AGR: Age Group fairness Reward for Bias Mitigation in LLMs

  • 构建针对年龄偏见的偏好与指令微调数据集,支持强化学习优化
  • 引入年龄公平奖励(AGR)后,跨年龄组响应准确率显著提升,差异缩小40%以上
  • 适合关注大模型公平性、社会影响的研究者与开发者使用

大语言模型可能表现出年龄偏见,导致不同年龄群体受到不平等对待。尽管已有研究关注种族和性别偏见,但年龄偏见仍缺乏系统探索。由于缺乏用于年龄偏见的指令微调和偏好数据集,其检测与度量困难,现有微调方法也极少关注年龄相关公平性。本文构建了用于强化学习人类反馈(RLHF)的年龄偏见偏好数据集和指令微调数据集,并提出一种年龄公平奖励(AGR),旨在减少大模型在不同年龄群体间响应质量的差异。大量实验表明,该奖励显著提升了响应准确性,并降低了跨年龄组的表现差距。代码与数据集已公开于匿名链接。

原文摘要 · Abstract (English)

LLMs can exhibit age biases, resulting in unequal treatment of individuals across age groups. While much research has addressed racial and gender biases, age bias remains little explored. The scarcity of instruction-tuning and preference datasets for age bias hampers its detection and measurement, and existing fine-tuning methods seldom address age-related fairness. In this paper, we construct age bias preference datasets and instruction-tuning datasets for RLHF. We introduce ARG, an age fairness reward to reduce differences in the response quality of LLMs across different age groups. Extensive experiments demonstrate that this reward significantly improves response accuracy and reduces performance disparities across age groups. Our source code and datasets are available at the anonymous \href{https://anonymous.4open.science/r/FairRLHF-D445/readme.md}{link}.

大模型公平性年龄偏见强化学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。