让机器人通过少量示范自动生成可复用技能,边执行边学习改进。
Teach and Grow: An Agent-Centered Architecture for General Robot Learning
- 基于演示生成可闭环运行的技能模块,支持任务分解与组合。
- 在新场景中自主选择工具、调整策略,失败后自动修复路径。
- 适合需要持续学习和泛化能力的通用机器人系统研发者。
端到端视觉-语言-动作(VLA)和世界动作模型为通用机器人提供了简洁路径,但其可靠性受限于已验证的物理覆盖范围。当遇到未见过的物体、传感器、机体或接触情况且无可靠备用方案时,故障修正需新机器人数据、策略更新和回归测试,形成反复训练负担,称为‘重训税’。不同于文本数据,具身数据通常需通过实际机器操作生成。我们提出‘教与成长’学习(Teach-and-Grow Learning, TGL),一种以智能体为中心的通用机器人学习架构。在通用形式下,多模态智能体将少数成功示范转化为可复用的技能块——针对有意义子目标的闭环行为。在新场景中,智能体进行语义定位、组合这些技能块,选择已学或几何工具,观察物理结果,并在执行偏离意图时调整路径。技能库存储可执行行为,结构化经验记忆保留成功、失败及修复记录。新任务无需特定任务策略重训练。LIBERO评估达到当前最优性能;受控实验揭示技能归纳、持久复用与智能体主导的适应性。最后,我们提出‘教与成长’扩展定律假设:若X代表有效可复用经验,则未来任务误差与教学需求将随X按幂律趋近不可再降的下限。因此,该架构将部署视为持续学习阶段,使一个任务能令下一个任务更易完成。
原文摘要 · Abstract (English)
End-to-end vision-language-action (VLA) and world-action models offer an elegant route to general-purpose robotics, but their reliability is bounded by validated physical coverage. When an unfamiliar object, sensor, embodiment, or contact falls outside that coverage and no validated fallback exists, correcting the failure requires new robot data, a policy update, and regression testing. This recurring burden is the retraining tax. Unlike text, embodied data must often be created by operating machines. We present Teach-and-Grow Learning (TGL), an agent-centered architecture for general robot learning. In its general form, a multimodal agent turns a few successful demonstrations into reusable Skill Blocks: closed-loop behaviors for meaningful subgoals. In a new scene, the agent grounds and composes these blocks, selects learned or geometric tools, observes the physical outcome, and revises the route when execution departs from intent. A Skill Library stores executable behavior, while structured Experience Memory carries forward success, failure, and repair. New tasks are acquired without task-specific policy retraining. Our LIBERO evaluation attains state-of-the-art performance; controlled studies expose skill induction, persistent reuse, and agent-directed adaptation. Finally, we propose the Teach-and-Grow scaling-law hypothesis: if X denotes effective reusable experience, future-task error and teaching demand should approach irreducible floors as power laws in X. The architecture therefore treats deployment as a period of continued learning, in which one task can make the next easier.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。