arXiv:2504.13972cs.CYcs.AI2025-04被引 1

评估者理性影响大模型对齐效果,高理性者反馈更稳定

Governance Challenges in Reinforcement Learning from Human Feedback: Evaluator Rationality and Reinforcement Stability

  • 通过对比高/低理性评估者,发现理性程度影响反馈一致性
  • 低理性评估者决策差异显著(p<0.01),反馈波动大
  • 建议筛选评估者并加权聚合反馈,提升对齐可靠性

基于人类反馈的强化学习(RLHF)是使大型语言模型(LLMs)与人类价值观保持一致的核心方法。然而该过程仍面临评估者偏见、不一致及反馈不可靠等治理挑战。本研究考察了评估者认知能力,特别是理性水平,对强化信号稳定性的影响。一项对照实验比较了高理性与低理性参与者,结果表明高理性评估者生成的反馈在一致性与专家对齐度上显著更优。相比之下,低理性参与者表现出显著的反馈决策变异性(p < 0.01)。为改善RLHF治理,本文建议实施评估者预筛选、系统性反馈一致性审计以及可靠性加权的强化信号聚合。这些措施可增强人工智能对齐流程的公平性、透明性与鲁棒性。

原文摘要 · Abstract (English)

Reinforcement Learning from Human Feedback (RLHF) is central in aligning large language models (LLMs) with human values and expectations. However, the process remains susceptible to governance challenges, including evaluator bias, inconsistency, and the unreliability of feedback. This study examines how the cognitive capacity of evaluators, specifically their level of rationality, affects the stability of reinforcement signals. A controlled experiment comparing high-rationality and low-rationality participants reveals that evaluators with higher rationality scores produce significantly more consistent and expert-aligned feedback. In contrast, lower-rationality participants demonstrate considerable variability in their reinforcement decisions ($p < 0.01$). To address these challenges and improve RLHF governance, we recommend implementing evaluator pre-screening, systematic auditing of feedback consistency, and reliability-weighted reinforcement aggregation. These measures enhance the fairness, transparency, and robustness of AI alignment pipelines.

RLHF评估者对齐

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。