arXiv:2503.09025cs.CL2025-03NAACL被引 11

RLHF难以消除语言模型对非裔美国人的隐性偏见,反而可能固化已有偏见。

Aligning to What? Limits to RLHF Based Alignment

  • 对比DPO、ORPO等方法,测试其对模型隐性与显性偏见的影响
  • SFT预训练会固化偏见,使后续RLHF效果减弱
  • 适用于研究模型偏见的学者及负责任的AI开发者

强化学习从人类反馈(RLHF)被广泛用于对齐大语言模型(LLMs)与人类偏好。然而,其在缓解深层偏见方面的有效性仍不明确。本研究探讨了RLHF与LLM中隐性及显性偏见的关系,特别关注对非裔美国人的偏见。我们采用DPO、ORPO和RLOO等技术对Llama 3 8B进行微调,并通过匹配语调探测和显性偏见测试评估模型的偏见水平。在不同基础模型与数据集上对DPO进行额外测试,发现:在RLHF前进行SFT会固化模型偏见。此外,我们将偏见测量工具扩展至多模态模型。实验表明,当前对齐技术难以有效处理如隐性偏见等模糊任务,亟需更高质量的数据集、数据构建方法或对齐工具。

原文摘要 · Abstract (English)

Reinforcement Learning from Human Feedback (RLHF) is increasingly used to align large language models (LLMs) with human preferences. However, the effectiveness of RLHF in addressing underlying biases remains unclear. This study investigates the relationship between RLHF and both covert and overt biases in LLMs, particularly focusing on biases against African Americans. We applied various RLHF techniques (DPO, ORPO, and RLOO) to Llama 3 8B and evaluated the covert and overt biases of the resulting models using matched-guise probing and explicit bias testing. We performed additional tests with DPO on different base models and datasets; among several implications, we found that SFT before RLHF calcifies model biases. Additionally, we extend the tools for measuring biases to multi-modal models. Through our experiments we collect evidence that indicates that current alignment techniques are inadequate for nebulous tasks such as mitigating covert biases, highlighting the need for capable datasets, data curating techniques, or alignment tools.

RLHF模型偏见偏见检测对齐

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。