大行为模型在多任务灵巧操作中表现更优,预训练越充分越强。
A Careful Examination of Large Behavior Models for Multitask Dexterous Manipulation

- 扩展扩散策略,用仿真与真实数据训练多任务机器人模型
- 多任务预训练显著提升成功率和鲁棒性,新任务学习仅需少量数据
- 性能随预训练规模与多样性增长而可预测提升,适合研究通用机器人智能
近年来,机器人灵巧操作取得显著进展,模仿学习策略成功执行了许多难以建模的复杂任务。与此同时,数据与模型规模的扩大催生了强大的语言与视觉基础模型,推动构建通用机器人基础模型的大型项目。尽管这些模型备受关注和投资,但其实验室外的真实性能评估仍具挑战性,制约了发展速度并阻碍对当前能力的深入理解。本文通过扩展扩散策略,系统评估多任务机器人操作策略——大行为模型(LBMs),涵盖仿真与真实数据。提出并验证了一套评估流程,实现统计置信度下的能力分析。在受控环境中,通过盲测随机试验对比单任务基线,结果表明:多任务预训练显著提升策略成功率与鲁棒性,并能以极小数据量快速学会复杂新任务;性能随预训练规模与多样性增长而可预测提升。
原文摘要 · Abstract (English)
Robot manipulation has seen tremendous progress in recent years, with imitation learning policies enabling successful performance of dexterous and hard-to-model tasks. Concurrently, scaling data and model size has led to the development of capable language and vision foundation models, motivating large-scale efforts to create general-purpose robot foundation models. While these models have garnered significant enthusiasm and investment, meaningful evaluation of real-world performance remains a challenge, limiting both the pace of development and inhibiting a nuanced understanding of current capabilities. In this paper, we rigorously evaluate multitask robot manipulation policies, referred to as Large Behavior Models (LBMs), by extending the Diffusion Policy paradigm across a corpus of simulated and real-world robot data. We propose and validate an evaluation pipeline to rigorously analyze the capabilities of these models with statistical confidence. We compare against single-task baselines through blind, randomized trials in a controlled setting, using both simulation and real-world experiments. We find that multi-task pretraining makes the policies more successful and robust, and enables teaching complex new tasks more quickly, using a fraction of the data when compared to single-task baselines. Moreover, performance predictably increases as pretraining scale and diversity grows. Project page: https://toyotaresearchinstitute.github.io/lbm1/
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。