研究发现人类反馈训练让大模型过度使用特定词汇,存在对齐偏差。
Word Overuse and Alignment in Large Language Models: The Influence of Learning from Human Feedback
- 通过模拟人类反馈训练,识别出模型偏好的高频词汇
- 实验显示人类偏好包含特定词的文本,证实反馈导致词汇过用
- 揭示标注员与用户对词汇期待的差异,提醒对齐透明性重要
大型语言模型常过度使用如"delve"和"intricate"等词汇,但其成因尚不明确。本研究以Meta的Llama模型为基础,探究学习人类反馈(LHF)的影响,涵盖基于人类反馈的强化学习与直接偏好优化。提出一种简单方法检测可能由LHF引发的词汇偏好。通过模拟LHF流程并进行实验,发现参与者系统性偏好包含特定词汇的文本,从而更有力地将词汇过用归因于LHF。这种现象可视为一种对齐偏差,且研究揭示了标注员与模型使用者在词汇期望上的潜在分歧。工作推动可解释人工智能研究,强调对齐研究中数据与过程透明的重要性。
原文摘要 · Abstract (English)
Large Language Models (LLMs) are known to overuse certain terms like "delve" and "intricate." The exact reasons for these lexical choices, however, have been unclear. Using Meta's Llama model, this study investigates the contribution of Learning from Human Feedback (LHF), under which we subsume Reinforcement Learning from Human Feedback and Direct Preference Optimization. We present a straightforward procedure for detecting the lexical preferences of LLMs that are potentially LHF-induced. Next, we more conclusively link LHF to lexical overuse by experimentally emulating the LHF procedure and demonstrating that participants systematically prefer text variants that include certain words. This lexical overuse can be seen as a sort of misalignment, though our study highlights the potential divergence between the lexical expectations of different populations -- namely LHF workers versus LLM users. Our work contributes to the growing body of research on explainable artificial intelligence and emphasizes the importance of both data and procedural transparency in alignment research.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。