通过模仿莱维飞行,让机器人多任务学习更高效地探索环境。
Leveraging Temporally Extended Behavior Sharing for Multi-task Reinforcement Learning
- 用相关任务的策略引导探索,动态调整探索强度。
- 在复杂机器人场景中实现更高状态空间覆盖,提升样本效率。
- 适合需要高效数据利用的机器人多任务学习研究者。
多任务强化学习(MTRL)通过在多个任务间训练智能体,促进知识共享,从而提升样本效率和泛化能力。然而,由于收集多样化任务数据成本高昂,将MTRL应用于机器人领域仍具挑战。为此,我们提出MT-Lévy,一种新颖的探索策略,结合任务间行为共享与受莱维飞行启发的时序扩展探索,以增强MTRL环境中的样本效率。MT-Lévy利用相关任务训练的策略引导探索至关键状态,并根据任务成功率动态调整探索水平。该方法在复杂机器人环境中实现了更高效的全状态空间覆盖。实证结果表明,MT-Lévy显著提升了探索效率与样本利用率,消融实验进一步验证了各组件的贡献,证明结合行为共享与自适应探索策略能显著提升MTRL在机器人应用中的实用性。
原文摘要 · Abstract (English)
Multi-task reinforcement learning (MTRL) offers a promising approach to improve sample efficiency and generalization by training agents across multiple tasks, enabling knowledge sharing between them. However, applying MTRL to robotics remains challenging due to the high cost of collecting diverse task data. To address this, we propose MT-Lévy, a novel exploration strategy that enhances sample efficiency in MTRL environments by combining behavior sharing across tasks with temporally extended exploration inspired by Lévy flight. MT-Lévy leverages policies trained on related tasks to guide exploration towards key states, while dynamically adjusting exploration levels based on task success ratios. This approach enables more efficient state-space coverage, even in complex robotics environments. Empirical results demonstrate that MT-Lévy significantly improves exploration and sample efficiency, supported by quantitative and qualitative analyses. Ablation studies further highlight the contribution of each component, showing that combining behavior sharing with adaptive exploration strategies can significantly improve the practicality of MTRL in robotics applications.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。