arXiv:2410.06347cs.ROcs.AI2024-10被引 4

用预收集数据训练机器人完成多目标任务,无需在线试错。

Goal-Conditioned Decision Transformer for Multi-Goal Offline Reinforcement Learning

  • 将目标状态融入序列建模,实现离线多目标决策
  • 在复杂任务中超越在线基线,稀疏奖励下仍表现稳健
  • 适合缺乏在线交互资源的机器人研发团队

机器人强化学习面临样本效率低和跨目标泛化差的问题。尽管离线强化学习减少了对昂贵在线交互的需求,但其与目标条件策略及基于Transformer的架构结合仍研究不足。本文提出一种适用于离线多目标机器人的目标条件决策变换器,通过显式将目标状态纳入序列建模框架,仅利用预收集数据即可高效解决多种任务。我们在新发布的Franka Emika Panda平台离线数据集上验证该方法,实验表明,该方法在复杂任务中优于现有在线基线,并在稀疏奖励设置下保持鲁棒性,即使专家示范数据有限也表现良好。

原文摘要 · Abstract (English)

Reinforcement learning (RL) in robotics faces significant hurdles regarding sample efficiency and generalization across varying goals. While Offline RL mitigates the need for costly online interactions, its integration with goal-conditioned policies and transformer-based architectures remains underexplored. We introduce a Goal-Conditioned Decision Transformer adapted for offline multi-goal robotics. By explicitly incorporating goal states into the sequence modeling framework, our approach efficiently solves varying tasks using only pre-collected data. We validate this method on a newly released offline dataset for the Franka Emika Panda platform. Experimental results demonstrate that our approach outperforms state-of-the-art online baselines in complex tasks and maintains robustness in sparse-reward settings, even with limited expert demonstrations.

强化学习离线学习机器人Transformer

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。