厘清人类反馈在强化学习中的三种角色,指导更透明的模型对齐设计。
Three Models of RLHF Annotation: Extension, Evidence, and Authority
- 提出人类反馈的三种概念模型:延伸、证据与权威。
- 指出混淆三者会导致对齐失败,需针对不同任务选择合适模型。
- 建议将标注拆解为独立维度,分别适配对应模型。
基于偏好的对齐方法,尤其是人类反馈强化学习(RLHF),依赖人工标注者的判断来塑造大语言模型的行为。然而,这些判断的规范性角色通常未被明确。本文区分了三种概念模型:第一种是延伸——标注者扩展系统设计者自身对输出应然性的判断;第二种是证据——标注者提供关于道德、社会等事实的独立证据;第三种是权威——标注者作为更广泛群体代表,具有独立决定系统输出的权力。本文认为这些模型对如何征集、验证和聚合标注具有重要影响。通过回顾领域内里程碑论文,展示了它们隐含地运用了这些模型,并揭示了因无意或有意混淆模型而导致的失效模式,进而提出选择模型的规范性标准。核心建议是:RLHF流程设计者应将标注分解为可分离的维度,并为每个维度匹配最合适的模型,而非追求单一统一的流程。
原文摘要 · Abstract (English)
Preference-based alignment methods, most prominently Reinforcement Learning with Human Feedback (RLHF), use the judgments of human annotators to shape large language model behaviour. However, the normative role of these judgments is rarely made explicit. I distinguish three conceptual models of that role. The first is extension: annotators extend the system designers' own judgments about what outputs should be. The second is evidence: annotators provide independent evidence about some facts, whether moral, social or otherwise. The third is authority: annotators have some independent authority (as representatives of the broader population) to determine system outputs. I argue that these models have implications for how RLHF pipelines should solicit, validate and aggregate annotations. I survey landmark papers in the literature on RLHF and related methods to illustrate how they implicitly draw on these models, describe failure modes that come from unintentionally or intentionally conflating them, and offer normative criteria for choosing among them. My central recommendation is that RLHF pipeline designers should decompose annotation into separable dimensions and tailor each pipeline to the model most appropriate for that dimension, rather than seeking a single unified pipeline.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。