让机器人舰队通过视觉世界模型自主纠错,减少人工干预。
Multi-Task Interactive Robot Fleet Learning with Visual World Models
- 用视觉世界模型预测动作后果,实时监控异常。
- 异常检测器随自主能力提升自动调整标准,降低人工介入频率。
- 在仿真与真实场景中验证,适配多任务机器人部署。
大规模多任务机器人学习的进展为家庭和工业场景中部署机器人舰队提供了可能,使其能在多种环境中执行多样化任务。然而,面对现实世界的复杂性和不确定性,当前智能机器人常面临泛化性与鲁棒性不足的问题。我们提出Sirius-Fleet——一种多任务交互式机器人舰队学习框架,可在部署过程中监控机器人表现,并在必要时引入人类纠正动作。该框架采用视觉世界模型预测未来动作的结果,并构建异常预测器以判断动作是否可能导致异常。随着机器人自主能力的提升,异常预测器会自动调整其判断标准,从而逐步减少对人类干预的需求,减轻长期人力负担。在大型基准测试上,Sirius-Fleet展现出显著提升的多任务策略性能与监控准确性。我们在模拟环境中的RoboCasa和真实世界中的Mutex两个多样且大规模的多任务基准上验证了其有效性。更多信息见项目官网:https://ut-austin-rpl.github.io/sirius-fleet
原文摘要 · Abstract (English)
Recent advancements in large-scale multi-task robot learning offer the potential for deploying robot fleets in household and industrial settings, enabling them to perform diverse tasks across various environments. However, AI-enabled robots often face challenges with generalization and robustness when exposed to real-world variability and uncertainty. We introduce Sirius-Fleet, a multi-task interactive robot fleet learning framework to address these challenges. Sirius-Fleet monitors robot performance during deployment and involves humans to correct the robot's actions when necessary. We employ a visual world model to predict the outcomes of future actions and build anomaly predictors to predict whether they will likely result in anomalies. As the robot autonomy improves, the anomaly predictors automatically adapt their prediction criteria, leading to fewer requests for human intervention and gradually reducing human workload over time. Evaluations on large-scale benchmarks demonstrate Sirius-Fleet's effectiveness in improving multi-task policy performance and monitoring accuracy. We demonstrate Sirius-Fleet's performance in both RoboCasa in simulation and Mutex in the real world, two diverse, large-scale multi-task benchmarks. More information is available on the project website: https://ut-austin-rpl.github.io/sirius-fleet
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。