构建365项日常任务的仿真环境,助力通用机器人训练与评估
RoboCasa365: A Large-Scale Simulation Framework for Training and Benchmarking Generalist Robots
- 基于RoboCasa平台构建2500个多样化厨房场景
- 含超600小时真人演示与1600小时合成数据
- 支持多任务学习、基础模型训练等系统性评估
近期机器人学习进展推动了通用机器人在人类环境中执行日常任务的发展。然而,当前仍难以评估离这一目标有多远。该领域缺乏可复现的大规模基准测试。为此,我们提出RoboCasa365,一个面向家庭移动操作的综合性仿真基准。依托RoboCasa平台,该基准包含365项日常任务,覆盖2,500个多样化的厨房环境,拥有超过600小时的人类示范数据和超过1600小时的合成生成示范数据,是研究通用策略最丰富、规模最大的资源之一。RoboCasa365旨在支持多种问题设置下的系统性评估,包括多任务学习、机器人基础模型训练及终身学习。我们在该基准上对前沿方法进行了广泛实验,分析了任务多样性、数据集规模与环境变化对泛化性能的影响。结果揭示了影响通用机器人表现的关键因素,并为未来研究提供指导。
原文摘要 · Abstract (English)
Recent advances in robot learning have accelerated progress toward generalist robots that can perform everyday tasks in human environments. Yet it remains difficult to gauge how close we are to this vision. The field lacks a reproducible, large-scale benchmark for systematic evaluation. To fill this gap, we present RoboCasa365, a comprehensive simulation benchmark for household mobile manipulation. Built on the RoboCasa platform, RoboCasa365 introduces 365 everyday tasks across 2,500 diverse kitchen environments, with over 600 hours of human demonstration data and over 1600 hours of synthetically generated demonstration data -- making it one of the most diverse and large-scale resources for studying generalist policies. RoboCasa365 is designed to support systematic evaluations for different problem settings, including multi-task learning, robot foundation model training, and lifelong learning. We conduct extensive experiments on this benchmark with state-of-the-art methods and analyze the impacts of task diversity, dataset scale, and environment variation on generalization. Our results provide new insights into what factors most strongly affect the performance of generalist robots and inform strategies for future progress in the field.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。