通过时间加权对比学习,从成功与失败演示中高效提取密集奖励信号。
TW-CRL: Time-Weighted Contrastive Reward Learning for Efficient Inverse Reinforcement Learning

- 利用成功与失败轨迹中的时序信息,构建对比学习目标
- 在导航与机器人操作任务中显著提升学习效率与鲁棒性
- 特别适合存在隐蔽陷阱状态的复杂决策场景
强化学习中的片段式任务常因奖励稀疏和高维状态空间导致学习效率低下。此外,这些任务常包含隐性“陷阱状态”——不可逆的失败状态,虽阻止任务完成,但不提供明确负向奖励,难以引导智能体规避重复错误。为此,我们提出时间加权对比奖励学习(TW-CRL),一种逆强化学习框架,同时利用成功与失败的示范数据。通过引入时序信息,TW-CRL学习一个密集奖励函数,识别与成功或失败相关的关键状态。该方法不仅使智能体能避开陷阱状态,还鼓励其进行超越简单模仿专家轨迹的有意义探索。在导航任务与机器人操作基准上的实证评估表明,TW-CRL优于现有最先进方法,实现了更高的学习效率与鲁棒性。
原文摘要 · Abstract (English)
Episodic tasks in Reinforcement Learning (RL) often pose challenges due to sparse reward signals and high-dimensional state spaces, which hinder efficient learning. Additionally, these tasks often feature hidden "trap states" -- irreversible failures that prevent task completion but do not provide explicit negative rewards to guide agents away from repeated errors. To address these issues, we propose Time-Weighted Contrastive Reward Learning (TW-CRL), an Inverse Reinforcement Learning (IRL) framework that leverages both successful and failed demonstrations. By incorporating temporal information, TW-CRL learns a dense reward function that identifies critical states associated with success or failure. This approach not only enables agents to avoid trap states but also encourages meaningful exploration beyond simple imitation of expert trajectories. Empirical evaluations on navigation tasks and robotic manipulation benchmarks demonstrate that TW-CRL surpasses state-of-the-art methods, achieving improved efficiency and robustness.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。