arXiv:2410.05026cs.LGcs.RO2024-10ICML被引 9

主动选择演示任务,用最少数据高效训练多任务机器人策略。

Active Fine-Tuning of Multi-Task Policies

  • 智能挑选最值得演示的任务,提升学习效率。
  • 在有限演示次数下,显著提升多任务策略性能。
  • 适合需要快速适应多个新任务的机器人系统。

预训练通用策略在机器人学习中日益重要,因其能快速适应新任务。现有方法通常依赖为特定任务收集新示范并使用模仿学习(如行为克隆)。但当需学习多个任务时,如何决定演示哪些任务及频率成为关键问题。本文研究此多任务挑战,提出交互式框架,让代理自适应选择需演示的任务。提出AMF(主动多任务微调)算法,在有限示范预算下,通过收集信息增益最大的示范来最大化多任务策略性能。在合理假设下推导了性能保证,并在复杂高维环境中实证其有效性,能高效微调神经策略。

原文摘要 · Abstract (English)

Pre-trained generalist policies are rapidly gaining relevance in robot learning due to their promise of fast adaptation to novel, in-domain tasks. This adaptation often relies on collecting new demonstrations for a specific task of interest and applying imitation learning algorithms, such as behavioral cloning. However, as soon as several tasks need to be learned, we must decide which tasks should be demonstrated and how often? We study this multi-task problem and explore an interactive framework in which the agent adaptively selects the tasks to be demonstrated. We propose AMF (Active Multi-task Fine-tuning), an algorithm to maximize multi-task policy performance under a limited demonstration budget by collecting demonstrations yielding the largest information gain on the expert policy. We derive performance guarantees for AMF under regularity assumptions and demonstrate its empirical effectiveness to efficiently fine-tune neural policies in complex and high-dimensional environments.

机器人学习多任务主动学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。