arXiv:2507.08707cs.LGcs.RO2025-07被引 1

从低质量演示中高效学习长时程对抗任务的强化学习新方法

SPLASH! Sample-efficient Preference-based inverse reinforcement learning for Long-horizon Adversarial tasks from Suboptimal Hierarchical demonstrations

  • 基于偏好学习,从不完美的分层示范中提取奖励信号
  • 在仿真海事抓旗任务中显著超越现有方法,实现实时迁移
  • 适合需长周期决策与对抗策略的机器人场景

逆强化学习(IRL)为从人类示范中学习复杂机器人任务提供了强大范式。然而,多数方法假设存在专家示范,这在实际中常不成立。部分允许示范次优的方法无法处理长时程目标或对抗性任务。许多理想的机器人能力同时具备这两类特征,凸显了IRL在生成可部署机器人代理方面的关键短板。本文提出样本高效的基于偏好的逆强化学习方法SPLASH,用于从次优分层示范中学习长时程对抗任务,显著推进了该领域前沿。我们在仿真海事抓旗任务中对SPLASH进行了实证验证,并通过模拟到现实的迁移实验展示了其在自主无人水面艇上的实际应用潜力。结果表明,所提方法在从次优示范中学习奖励函数方面显著优于当前最优方法。

原文摘要 · Abstract (English)

Inverse Reinforcement Learning (IRL) presents a powerful paradigm for learning complex robotic tasks from human demonstrations. However, most approaches make the assumption that expert demonstrations are available, which is often not the case. Those that allow for suboptimality in the demonstrations are not designed for long-horizon goals or adversarial tasks. Many desirable robot capabilities fall into one or both of these categories, thus highlighting a critical shortcoming in the ability of IRL to produce field-ready robotic agents. We introduce Sample-efficient Preference-based inverse reinforcement learning for Long-horizon Adversarial tasks from Suboptimal Hierarchical demonstrations (SPLASH), which advances the state-of-the-art in learning from suboptimal demonstrations to long-horizon and adversarial settings. We empirically validate SPLASH on a maritime capture-the-flag task in simulation, and demonstrate real-world applicability with sim-to-real translation experiments on autonomous unmanned surface vehicles. We show that our proposed methods allow SPLASH to significantly outperform the state-of-the-art in reward learning from suboptimal demonstrations.

逆强化学习长时程任务对抗任务机器人控制

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。