arXiv:2607.11432cs.LG2026-07

让强化学习理解人类对轨迹的‘无法比较’判断,提升奖励建模的合理性。

Generalizing Preference-based Reinforcement Learning: a Rationality Model for Incomparability

论文配图:Generalizing Preference-based Reinforcement Learning: a Rationality Model for Incomparability
图 1 · 摘自论文原文
  • 基于人类对轨迹对的‘不可比’反馈,构建多维奖励函数
  • 在模拟环境中准确重建专家偏好并恢复最优策略边界
  • 适用于专家理性不一致或偏好模糊的场景

本文研究从人类专家提供的轨迹对比较中学习强化学习策略的问题。通过形式化专家可标记轨迹对为‘不可比’(即无优劣之分)的新设定,扩展了偏好强化学习。提出一种受Bradley-Terry启发的理性模型,能有效捕捉不可比性并推断多维奖励函数,分析其性质,并提供参数学习的样本复杂度理论。在模拟环境中验证该模型可准确重构与专家偏好一致的奖励函数,恢复帕累托前沿策略,并对不同水平的专家理性表现出良好鲁棒性。

原文摘要 · Abstract (English)

In this work, we study the reinforcement learning (RL) problem from pairwise trajectory comparisons provided by a human expert. We generalize preference-based RL by formalizing a novel setting in which the expert can also label trajectory pairs as incomparable, i.e., when neither trajectory dominates the other. We introduce the learning problem and the desiderata that its solution should satisfy. Then, we propose a novel Bradley-Terry-inspired rationality model that effectively captures incomparabilities and infers a multi-dimensional reward function, and we study its properties. We provide a sample complexity analysis for learning the model parameters when a dataset is available. Finally, we evaluate our model's ability to reconstruct a reward function that aligns with the expert's comparisons in simulated environments and to recover the Pareto frontier of policies, along with a robustness analysis across varying levels of expert rationality.

偏好学习强化学习多维奖励不可比

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。