arXiv:2502.02921cs.LG2025-02ICML被引 3

通过批量切割假设空间,提升强化学习中奖励对齐的鲁棒性。

Robust Reward Alignment via Hypothesis Space Batch Cutting

  • 基于人类偏好批次的几何切割机制,逐步缩小奖励假设空间。
  • 在错误偏好占比高达50%时仍保持性能优势,优于现有方法。
  • 无需识别错误偏好,适合实际应用中人类反馈不可靠的场景。

强化学习与最优控制中的奖励设计极具挑战性。基于偏好的对齐方法通过人类提供的轨迹对排序来学习奖励,但现有方法常因未知的错误人类偏好而缺乏鲁棒性。本文提出一种基于新视角——假设空间批量切割的鲁棒高效奖励对齐方法。该方法通过基于批次人类偏好的‘切割’迭代优化奖励假设空间。每批次内,基于分歧查询的人类偏好通过投票函数分组以确定合适切割,确保人类查询复杂度有界。为应对未知错误偏好,引入保守切割策略,防止错误偏好造成过激切割,从而保证对错误偏好的可证明鲁棒性,且无需显式识别错误偏好。我们在多种任务的模型预测控制设置下评估该方法。结果表明,在无错误设置中性能与最先进方法相当或更优;而在高比例错误偏好情况下,显著优于现有方法。

原文摘要 · Abstract (English)

Reward design in reinforcement learning and optimal control is challenging. Preference-based alignment addresses this by enabling agents to learn rewards from ranked trajectory pairs provided by humans. However, existing methods often struggle from poor robustness to unknown false human preferences. In this work, we propose a robust and efficient reward alignment method based on a novel and geometrically interpretable perspective: hypothesis space batched cutting. Our method iteratively refines the reward hypothesis space through "cuts" based on batches of human preferences. Within each batch, human preferences, queried based on disagreement, are grouped using a voting function to determine the appropriate cut, ensuring a bounded human query complexity. To handle unknown erroneous preferences, we introduce a conservative cutting method within each batch, preventing erroneous human preferences from making overly aggressive cuts to the hypothesis space. This guarantees provable robustness against false preferences, while eliminating the need to explicitly identify them. We evaluate our method in a model predictive control setting across diverse tasks. The results demonstrate that our framework achieves comparable or superior performance to state-of-the-art methods in error-free settings while significantly outperforming existing methods when handling a high percentage of erroneous human preferences.

强化学习奖励对齐鲁棒性人类偏好

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。