arXiv:2504.16054cs.LGcs.RO2025-04被引 1.7k

让机器人在陌生家庭中自主完成清洁等复杂任务

$π_{0.5}$: a Vision-Language-Action Model with Open-World Generalization

论文配图:$π_{0.5}$: a Vision-Language-Action Model with Open-World Generalization
图 1 · 摘自论文原文
  • 通过多源数据联合训练提升模型泛化能力
  • 首次实现端到端系统在新环境中完成长时序精细操作
  • 适合研究机器人通用智能与真实世界应用的学者

为了让机器人真正有用,必须使其能在实验室外的真实世界中执行实际任务。尽管视觉-语言-动作(VLA)模型在端到端机器人控制方面表现优异,但其在开放环境中的泛化能力仍不明确。本文提出 $π_{0.5}$,基于 $π_{0}$ 模型,通过异构任务的联合训练实现广泛泛化。该模型融合多机器人数据、高层语义预测、网络数据等多种来源,支持广泛可泛化的现实世界机器人操作。系统采用联合训练与混合多模态样本策略,结合图像观测、语言指令、物体检测、语义子任务预测和低层动作。实验表明,此类知识迁移对有效泛化至关重要,并首次展示端到端学习系统可在完全陌生的住宅中完成如清理厨房或卧室等长周期、高精度操作。

原文摘要 · Abstract (English)

In order for robots to be useful, they must perform practically relevant tasks in the real world, outside of the lab. While vision-language-action (VLA) models have demonstrated impressive results for end-to-end robot control, it remains an open question how far such models can generalize in the wild. We describe $π_{0.5}$, a new model based on $π_{0}$ that uses co-training on heterogeneous tasks to enable broad generalization. $π_{0.5}$\ uses data from multiple robots, high-level semantic prediction, web data, and other sources to enable broadly generalizable real-world robotic manipulation. Our system uses a combination of co-training and hybrid multi-modal examples that combine image observations, language commands, object detections, semantic subtask prediction, and low-level actions. Our experiments show that this kind of knowledge transfer is essential for effective generalization, and we demonstrate for the first time that an end-to-end learning-enabled robotic system can perform long-horizon and dexterous manipulation skills, such as cleaning a kitchen or bedroom, in entirely new homes.

机器人控制多模态学习泛化能力端到端

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。