优化微调策略让通用机器人抓取模型在少样本下表现更优
Effective Tuning Strategies for Generalist Robot Manipulation Policies
- 系统测试2500次实验,找出微调关键因素
- 少样本下性能超越现有模仿学习算法
- 为通用机器人策略微调提供实用指南
通用机器人操作策略(GMPs)具备跨任务、设备和环境泛化潜力,但受限于难以收集覆盖广泛领域的动作数据,在分布外场景中仍表现不佳。尽管微调能以少量样本快速适配新任务,我们发现其性能显著受微调策略设计影响。本文通过深入的实证研究,系统评估了动作空间、策略头、监督信号及可调参数等关键因素,单个配置需完成2500次仿真回滚。基于分析结果,我们总结出有效设计原则,并验证在低数据条件下,合理微调后的GMPs显著优于当前最优模仿学习方法。本工作为微调型GMPs研究建立新基准,为社区提供实用工具支持。
原文摘要 · Abstract (English)
Generalist robot manipulation policies (GMPs) have the potential to generalize across a wide range of tasks, devices, and environments. However, existing policies continue to struggle with out-of-distribution scenarios due to the inherent difficulty of collecting sufficient action data to cover extensively diverse domains. While fine-tuning offers a practical way to quickly adapt a GMPs to novel domains and tasks with limited samples, we observe that the performance of the resulting GMPs differs significantly with respect to the design choices of fine-tuning strategies. In this work, we first conduct an in-depth empirical study to investigate the effect of key factors in GMPs fine-tuning strategies, covering the action space, policy head, supervision signal and the choice of tunable parameters, where 2,500 rollouts are evaluated for a single configuration. We systematically discuss and summarize our findings and identify the key design choices, which we believe give a practical guideline for GMPs fine-tuning. We observe that in a low-data regime, with carefully chosen fine-tuning strategies, a GMPs significantly outperforms the state-of-the-art imitation learning algorithms. The results presented in this work establish a new baseline for future studies on fine-tuned GMPs, and provide a significant addition to the GMPs toolbox for the community.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。