用专家监督信号增强逆强化学习,提升复杂环境下的学习效率。
Hybrid-AIRL: Enhancing Inverse Reinforcement Learning with Supervised Expert Guidance
- 结合专家数据的监督损失与随机正则化,改进奖励函数推断
- 在德州扑克和多个基准测试中样本效率更高、训练更稳定
- 适合需要高可靠性的现实场景逆强化学习任务
对抗性逆强化学习(AIRL)在解决强化学习中的稀疏奖励问题上表现出潜力,通过专家示范推断密集奖励函数。然而,其在高度复杂、信息不完全的场景中的表现仍不明确。为此,我们在二人有限注额德州扑克(HULHE)这一具有稀疏延迟奖励和显著不确定性的领域评估了AIRL,发现其难以推断出足够信息的奖励函数。为此,我们提出Hybrid-AIRL(H-AIRL),通过引入基于专家数据的监督损失和随机正则化机制,增强奖励推断与策略学习。我们在精选的Gymnasium基准和HULHE扑克设置上评估了H-AIRL。进一步通过可视化分析学习到的奖励函数,深入理解学习过程。实验结果表明,相较于AIRL,H-AIRL在样本效率和学习稳定性方面均有显著提升,证明了将监督信号融入逆强化学习的有效性,并确立了其在挑战性真实场景中的应用前景。
原文摘要 · Abstract (English)
Adversarial Inverse Reinforcement Learning (AIRL) has shown promise in addressing the sparse reward problem in reinforcement learning (RL) by inferring dense reward functions from expert demonstrations. However, its performance in highly complex, imperfect-information settings remains largely unexplored. To explore this gap, we evaluate AIRL in the context of Heads-Up Limit Hold'em (HULHE) poker, a domain characterized by sparse, delayed rewards and significant uncertainty. In this setting, we find that AIRL struggles to infer a sufficiently informative reward function. To overcome this limitation, we contribute Hybrid-AIRL (H-AIRL), an extension that enhances reward inference and policy learning by incorporating a supervised loss derived from expert data and a stochastic regularization mechanism. We evaluate H-AIRL on a carefully selected set of Gymnasium benchmarks and the HULHE poker setting. Additionally, we analyze the learned reward function through visualization to gain deeper insights into the learning process. Our experimental results show that H-AIRL achieves higher sample efficiency and more stable learning compared to AIRL. This highlights the benefits of incorporating supervised signals into inverse RL and establishes H-AIRL as a promising framework for tackling challenging, real-world settings.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。