arXiv:2602.08092cs.AIcs.ET2026-02被引 2

在群体评价中,传统强化学习会因多数人迎合而误判目标,新方法通过评判评价者来源纠正偏差。

Objective Decoupling in Social Reinforcement Learning: Recovering Ground Truth from Sycophantic Majorities

  • 不依赖多数投票,而是评估反馈来源的可靠性
  • 在多数评价者合谋误导时仍能收敛到真实目标
  • 适合存在偏见或合谋评价的社会化强化学习场景

当前人工智能对齐策略基于一个脆弱假设:人类反馈虽有噪声,但本质是真实信号。本文指出这一假设为强化学习中的‘教条4’。我们证明,在静态环境中该假设成立,但在社交场景中,评价者可能谄媚、懒惰或恶意,导致标准强化学习代理出现‘目标解耦’——学习目标与潜在真实目标永久偏离,必然导致对齐失败。为此,我们提出认知源对齐(ESA)方法。不同于依赖统计共识的鲁棒方法,ESA通过稀疏安全公理判断反馈来源而非信号本身。理论证明,这种‘评判评价者’机制可保证即使多数评价者存在偏见,仍能收敛至真实目标。实证表明,传统共识方法在多数合谋时失效,而我们的方法成功恢复最优策略。

原文摘要 · Abstract (English)

Contemporary AI alignment strategies rely on a fragile premise: that human feedback, while noisy, remains a fundamentally truthful signal. In this paper, we identify this assumption as Dogma 4 of Reinforcement Learning (RL). We demonstrate that while this dogma holds in static environments, it fails in social settings where evaluators may be sycophantic, lazy, or adversarial. We prove that under Dogma 4, standard RL agents suffer from what we call Objective Decoupling, a structural failure mode where the agent's learned objective permanently separates from the latent ground truth, guaranteeing convergence to misalignment. To resolve this, we propose Epistemic Source Alignment (ESA). Unlike standard robust methods that rely on statistical consensus (trusting the majority), ESA utilizes sparse safety axioms to judge the source of the feedback rather than the signal itself. We prove that this "judging the judges" mechanism guarantees convergence to the true objective, even when a majority of evaluators are biased. Empirically, we show that while traditional consensus methods fail under majority collusion, our approach successfully recovers the optimal policy.

强化学习对齐问题社会性反馈

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。