MAML-TRPO在机器人操作任务中实现快速适应,但泛化能力仍有提升空间。
Evaluating Model-Agnostic Meta-Learning on MetaWorld ML10 Benchmark: Fast Adaptation in Robotic Manipulation Tasks
- 用MAML+TRPO从少量数据中快速学习通用初始化策略。
- 单次梯度更新后成功率达21.0%(训练)和13.2%(测试)。
- 不同操作技能适应效果差异大,0%到80%不等,适合研究高效泛化方法者参考。
元学习算法可在少量数据下实现对新任务的快速适应,这对真实世界机器人系统至关重要。本文在MetaWorld ML10基准上评估了模型无关元学习(MAML)与信任域策略优化(TRPO)的结合,该基准包含10个多样化的机器人操作任务。实验验证了MAML-TRPO学习通用初始化的能力,使其能在推、拾取、抽屉操作等语义不同的行为中实现少样本适应。结果表明,经过一次梯度更新后,模型在训练任务上的成功率达到21.0%,测试任务为13.2%。然而,在元训练过程中出现泛化差距:测试性能趋于平稳而训练性能持续上升。任务级分析显示适应效果差异显著,成功率在0%至80%之间波动。这些发现揭示了基于梯度的元学习在多样化机器人操作中的潜力与局限,并指出了未来在任务感知适应与结构化策略架构方面的研究方向。
原文摘要 · Abstract (English)
Meta-learning algorithms enable rapid adaptation to new tasks with minimal data, a critical capability for real-world robotic systems. This paper evaluates Model-Agnostic Meta-Learning (MAML) combined with Trust Region Policy Optimization (TRPO) on the MetaWorld ML10 benchmark, a challenging suite of ten diverse robotic manipulation tasks. We implement and analyze MAML-TRPO's ability to learn a universal initialization that facilitates few-shot adaptation across semantically different manipulation behaviors including pushing, picking, and drawer manipulation. Our experiments demonstrate that MAML achieves effective one-shot adaptation with clear performance improvements after a single gradient update, reaching final success rates of 21.0% on training tasks and 13.2% on held-out test tasks. However, we observe a generalization gap that emerges during meta-training, where performance on test tasks plateaus while training task performance continues to improve. Task-level analysis reveals high variance in adaptation effectiveness, with success rates ranging from 0% to 80% across different manipulation skills. These findings highlight both the promise and current limitations of gradient-based meta-learning for diverse robotic manipulation, and suggest directions for future work in task-aware adaptation and structured policy architectures.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。