arXiv:2604.15483cs.LGcs.RO2026-04被引 114

一个能零样本执行复杂任务的通用机器人模型,无需训练即可操作咖啡机。

$π_{0.7}$: a Steerable Generalist Robotic Foundation Model with Emergent Capabilities

论文配图:$π_{0.7}$: a Steerable Generalist Robotic Foundation Model with Emergent Capabilities
图 1 · 摘自论文原文
  • 通过多模态提示条件化训练,实现指令精准引导。
  • 零样本跨平台通用性,可完成折叠衣物等未见过的任务。
  • 适合需要强泛化能力的机器人研发与部署场景。

我们提出一种新型机器人基础模型 $π_{0.7}$,可在多种场景下实现开箱即用的强性能表现。该模型能在未见过的环境中理解多样化语言指令,包括涉及多种厨房电器的多阶段任务;具备零样本跨机体泛化能力,例如从未见过的折叠衣物任务也能完成;还可直接执行如操作咖啡机等高难度任务,性能媲美经过强化学习微调的专用模型。其核心思想是在训练中使用多样化的上下文条件化,这些条件信息(包含在提示中)不仅包括描述应执行动作的语言指令,还包含任务执行方式、策略及子目标图像等元数据。这使得 $π_{0.7}$ 能利用多样化数据源,包括示范数据、可能有误的自主生成数据(含失败案例)以及非机器人来源的数据。我们在多个机器人平台、多种任务上评估了 $π_{0.7}$,涵盖速度与灵巧性要求高的任务、语言跟随能力以及组合式任务泛化能力。

原文摘要 · Abstract (English)

We present a new robotic foundation model, called $π_{0.7}$, that can enable strong out-of-the-box performance in a wide range of scenarios. $π_{0.7}$ can follow diverse language instructions in unseen environments, including multi-stage tasks with various kitchen appliances, provide zero-shot cross-embodiment generalization, for example enabling a robot to fold laundry without seeing the task before, and perform challenging tasks such as operating an espresso machine out of the box at a level of performance that matches much more specialized RL-finetuned models. The main idea behind $π_{0.7}$ is to use diverse context conditioning during training. This conditioning information, contained in the prompt, makes it possible to steer the model precisely to perform many tasks with different strategies. It is conditioned not just on a language command that describes what it should do, but on additional multimodal information that also describes the manner or strategy in which it should do it, including metadata about task performance and subgoal images. This enables $π_{0.7}$ to use very diverse data, including demonstrations, potentially suboptimal (autonomous) data including failures, and data from non-robot sources. Our experiments evaluate $π_{0.7}$ across numerous tasks with multiple robot platforms, on tasks that require speed and dexterity, language following, and compositional task generalization.

机器人通用模型零样本多模态

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。