arXiv:2604.20074cs.LG2026-04IJCAI被引 21

利用无标签数据提升专家行为学习效率

Maximum Entropy Semi-Supervised Inverse Reinforcement Learning

  • 在最大熵逆强化学习中引入无监督轨迹的成对惩罚机制
  • 在高速驾驶和网格世界任务中性能优于传统方法
  • 适合有大量无标注数据但专家样本稀缺的场景

习得式学习(AL)常被建模为逆强化学习(IRL)问题。最大熵逆强化学习(MaxEnt-IRL)将最大熵原理融入IRL,解决了因存在大量可能匹配专家行为的策略而导致的歧义问题。本文研究了一种新设置:除专家轨迹外,还存在若干无监督轨迹。提出MESSI算法,将MaxEnt-IRL与半监督学习思想结合,通过轨迹间的成对惩罚机制融合无监督数据。在高速公路驾驶与网格世界任务中的实验表明,MESSI能有效利用无监督轨迹,显著提升MaxEnt-IRL性能。

原文摘要 · Abstract (English)

A popular approach to apprenticeship learning (AL) is to formulate it as an inverse reinforcement learning (IRL) problem. The MaxEnt-IRL algorithm successfully integrates the maximum entropy principle into IRL and unlike its predecessors, it resolves the ambiguity arising from the fact that a possibly large number of policies could match the expert's behavior. In this paper, we study an AL setting in which in addition to the expert's trajectories, a number of unsupervised trajectories is available. We introduce MESSI, a novel algorithm that combines MaxEnt-IRL with principles coming from semi-supervised learning. In particular, MESSI integrates the unsupervised data into the MaxEnt-IRL framework using a pairwise penalty on trajectories. Empirical results in a highway driving and grid-world problems indicate that MESSI is able to take advantage of the unsupervised trajectories and improve the performance of MaxEnt-IRL.

逆强化学习半监督最大熵

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。