从专家数据中自动选特征,构建可解释的奖励模型。
Learning Transparent Reward Models via Unsupervised Feature Selection
- 自动筛选状态特征,构造透明奖励模型。
- 在复杂机器人任务中逼近专家行为表现。
- 适合需要可解释奖励机制的研究者。
在机器人操作和自动驾驶等复杂现实任务中,获取专家示范往往比明确学习目标更容易。可通过行为克隆或逆强化学习来从专家数据中学习,后者允许使用分布外的数据进行训练,由推断出的奖励函数引导。本文提出一种新方法,通过自动选择的状态特征构建紧凑且透明的奖励模型。这些推断出的奖励具有显式形式,可直接用标准强化学习算法从零开始训练策略,使其紧密匹配专家行为。我们在多个具有连续高维状态空间的机器人环境中验证了该方法的有效性。
原文摘要 · Abstract (English)
In complex real-world tasks such as robotic manipulation and autonomous driving, collecting expert demonstrations is often more straightforward than specifying precise learning objectives and task descriptions. Learning from expert data can be achieved through behavioral cloning or by learning a reward function, i.e., inverse reinforcement learning. The latter allows for training with additional data outside the training distribution, guided by the inferred reward function. We propose a novel approach to construct compact and transparent reward models from automatically selected state features. These inferred rewards have an explicit form and enable the learning of policies that closely match expert behavior by training standard reinforcement learning algorithms from scratch. We validate our method's performance in various robotic environments with continuous and high-dimensional state spaces. Webpage: \url{https://sites.google.com/view/transparent-reward}.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。