arXiv:2501.01669cs.LGcs.RO2025-01中稿 · IJCAI被引 2

从不同任务中逆向学习通用奖励函数,实现跨场景迁移。

Inversely Learning Transferable Rewards via Abstracted States

  • 通过行为轨迹逆向学习抽象奖励函数,捕捉任务共性。
  • 在未见过的新任务实例中成功生成有效行为,验证可迁移性。
  • 适合需要快速适应新任务的机器人应用,如产线切换。

逆强化学习(IRL)在离散和连续域中已能准确从行为数据中学习底层奖励。下一步是学习内在偏好,使其在与观测任务相关但不同的场景中产生有用行为。在机器人应用中,这有助于将机器人集成到涉及新任务的生产线上,而无需重新编程。本文提出一种方法,从两个或更多不同实例的行为轨迹中逆向学习一个抽象奖励函数,并将其用于另一独立实例的任务行为学习。该步骤验证了其可迁移性与正确性。我们在 OpenAI Gym 和 AssistiveGym 的多个领域任务轨迹上进行了评估,结果表明所学抽象奖励函数可在相应领域未曾见过的新实例中成功学习任务行为。

原文摘要 · Abstract (English)

Inverse reinforcement learning (IRL) has progressed significantly toward accurately learning the underlying rewards in both discrete and continuous domains from behavior data. The next advance is to learn {\em intrinsic} preferences in ways that produce useful behavior in settings or tasks which are different but aligned with the observed ones. In the context of robotic applications, this helps integrate robots into processing lines involving new tasks (with shared intrinsic preferences) without programming from scratch. We introduce a method to inversely learn an abstract reward function from behavior trajectories in two or more differing instances of a domain. The abstract reward function is then used to learn task behavior in another separate instance of the domain. This step offers evidence of its transferability and validates its correctness. We evaluate the method on trajectories in tasks from multiple domains in OpenAI's Gym testbed and AssistiveGym and show that the learned abstract reward functions can successfully learn task behaviors in instances of the respective domains, which have not been seen previously.

逆强化学习奖励学习迁移能力机器人

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。