用可扩展的贝叶斯方法从专家行为中学习奖励函数,支持不确定性评估。
Q-based Variational Inverse Reinforcement Learning

- 通过变分优化最优Q值分布,推断奖励函数后验
- 在多个环境上实现高精度模仿学习,包括像素输入
- 首次支持从原始像素训练的贝叶斯逆强化学习
安全有益的AI系统需遵循人类偏好,但手动指定偏好往往不可行。逆强化学习(IRL)通过专家行为推断奖励函数来解决此问题。本文提出基于Q值的变分逆强化学习(QVIRL),一种新型贝叶斯IRL方法,主要通过学习最优Q值的变分分布,从专家演示中恢复奖励函数的后验分布。相比以往方法,QVIRL兼具可扩展性与不确定性量化能力,对安全关键场景和主动学习至关重要。我们在多种任务中验证了其性能,包括网格世界、月球着陆器、高速公路环境及两个ATARI游戏,覆盖静态专家数据和主动学习两种情形。QVIRL是首个能从原始像素观测进行训练的贝叶斯逆强化学习方法。
原文摘要 · Abstract (English)
The development of safe and beneficial AI requires that systems can learn and act in accordance with human preferences. However, explicitly specifying these preferences by hand is often infeasible. Inverse reinforcement learning (IRL) addresses this challenge by inferring preferences, represented as reward functions, from expert behaviour. We introduce Q-based Variational IRL (QVIRL), a novel Bayesian IRL method that recovers a posterior distribution over rewards from expert demonstrations via primarily learning a variational distribution over optimal Q-values. Unlike previous approaches, QVIRL combines scalability with uncertainty quantification, important for safety-critical applications as well as active learning. We demonstrate QVIRL's strong performance in apprenticeship learning across various tasks, including gridworlds, Lunar Lander, the Highway Environment, and two ATARI games both with static expert data and with active learning. It is the first method for Bayesian IRL that demonstrates training from raw pixel observations.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。