arXiv:2505.08619cs.LGcs.RO2025-05被引 4

用少量轨迹推断连续空间中的最优成本函数,速度快且精度高。

Cost Function Estimation Using Inverse Reinforcement Learning with Minimal Observations

  • 基于最大熵准则,迭代优化成本函数权重
  • 仅需少量观测即可收敛,比现有方法更快
  • 通过最优控制生成轨迹,信息量更丰富

我们提出一种迭代逆强化学习算法,用于在连续空间中推断最优成本函数。基于流行的最大熵准则,该方法迭代地寻找权重改进方向,并提出一种确定合适步长的方法,确保学习到的成本函数特征与示范轨迹特征保持一致。与类似方法相比,本算法可独立调节每条观测对分区函数的影响,无需大量样本,实现更快学习。我们通过求解最优控制问题生成样本轨迹,而非随机采样,从而获得更具信息量的轨迹。在多个模拟环境中,将本方法与两种前沿算法进行对比,验证了其优势。

原文摘要 · Abstract (English)

We present an iterative inverse reinforcement learning algorithm to infer optimal cost functions in continuous spaces. Based on a popular maximum entropy criteria, our approach iteratively finds a weight improvement step and proposes a method to find an appropriate step size that ensures learned cost function features remain similar to the demonstrated trajectory features. In contrast to similar approaches, our algorithm can individually tune the effectiveness of each observation for the partition function and does not need a large sample set, enabling faster learning. We generate sample trajectories by solving an optimal control problem instead of random sampling, leading to more informative trajectories. The performance of our method is compared to two state of the art algorithms to demonstrate its benefits in several simulated environments.

逆强化学习最优控制成本函数估计

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。