用专家轨迹训练模拟器,自动学奖励和环境,无需真实交互。
OffSim: Offline Simulator for Model-based Offline Inverse Reinforcement Learning
- 从专家轨迹中联合学习环境动态与奖励函数。
- 在MuJoCo上性能超越现有方法,提升探索与泛化能力。
- 适合想离线训练智能体的研究者,尤其数据有限时。
强化学习通常依赖预定义奖励函数的交互式模拟器进行策略训练。但构建模拟器和手动设定奖励函数往往耗时费力。为此,我们提出离线模拟器(OffSim),一种基于模型的离线逆强化学习框架,可直接从专家生成的状态-动作轨迹中模拟环境动力学与奖励结构。OffSim联合优化高熵转移模型与基于逆强化学习的奖励函数,以增强探索并提升学习奖励的泛化能力。利用这些学习到的组件,OffSim可在不与真实环境进一步交互的情况下训练策略。此外,我们引入OffSim$^+$,通过引入边际奖励机制,提升多数据集设置下的探索能力。大量MuJoCo实验表明,OffSim在性能上显著优于现有离线逆强化学习方法,验证了其有效性与鲁棒性。
原文摘要 · Abstract (English)
Reinforcement learning algorithms typically utilize an interactive simulator (i.e., environment) with a predefined reward function for policy training. Developing such simulators and manually defining reward functions, however, is often time-consuming and labor-intensive. To address this, we propose an Offline Simulator (OffSim), a novel model-based offline inverse reinforcement learning (IRL) framework, to emulate environmental dynamics and reward structure directly from expert-generated state-action trajectories. OffSim jointly optimizes a high-entropy transition model and an IRL-based reward function to enhance exploration and improve the generalizability of the learned reward. Leveraging these learned components, OffSim can subsequently train a policy offline without further interaction with the real environment. Additionally, we introduce OffSim$^+$, an extension that incorporates a marginal reward for multi-dataset settings to enhance exploration. Extensive MuJoCo experiments demonstrate that OffSim achieves substantial performance gains over existing offline IRL methods, confirming its efficacy and robustness.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。