让AI理解正在学习中的用户偏好,突破传统强化学习局限。
Learning the Preferences of a Learning Agent

- 将学习者建模为无遗憾或渐进最优的贝叶斯策略,推导偏好推理机制。
- 证明在特定学习动态下,可实现对奖励函数的稳定估计。
- 适合研究人机交互、自适应系统与在线学习场景的开发者参考。
为了让AI系统真正服务于人类,必须理解并遵循人类的价值与偏好。由于直接指定偏好极为困难,逆强化学习(IRL)旨在通过观察行为来推断偏好。然而,传统IRL假设人类行为近似最优,这在人类自身仍在学习如何最优行动时成为重大限制。本文形式化了学习型代理的偏好学习问题:一个观测者在在线观察学习者行为时,试图推断其正被(初始次优地)优化的潜在奖励函数。我们将学习者建模为无遗憾学习者或随时间收敛至最优贝叶斯策略的个体。在这些设定下,我们为多种偏好学习算法建立了理论保证,或证明此类保证在某些情况下不可能实现。
原文摘要 · Abstract (English)
For AI systems to be useful to humans, they must understand and act in accordance with our values and preferences. Since specifying preferences is a hard task, inverse reinforcement learning (IRL) aims to develop methods that allow for inferring preferences from observed behavior. However, IRL assumes the human to be approximately optimal. This is a big limitation in cases where the human themselves may be learning to act optimally in an environment. In this paper, we formalize the problem of learning the preferences of a learning agent: a predictor observes a learner acting online and tries to infer the underlying reward function being (initially suboptimally) optimized by the learner. We model the learner as either being no-regret, or as converging to an optimal Boltzmann policy over time. In each of these settings, we establish theoretical guarantees for various preference learning algorithms, or otherwise show that such guarantees are impossible.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。