arXiv:2608.01851cs.ROcs.AI2026-08综述

机器人学习正分两条路:固定权重或自动生成代码技能。

Weights or Skills? A Survey of Robot-Learning Techniques: from Action-Predicting Weights to Robots that Write their Own Skills

  • 按自我改进程度划分代码化策略,从零样本生成到持续进化闭环
  • 仅少数系统(如ASPIRE)实现执行反馈、记忆与搜索的开放循环
  • 揭示'技能'五种含义,仅代码技能可无梯度更新自进化

机器人学习正分裂为两种路径:将能力固化在冻结权重中的模型(如视觉-语言-动作模型),以及能自主编写和优化可执行代码技能的智能体。本综述以权重与技能为轴心组织领域,核心贡献是深入分析代码即策略方法的自我改进程度,从零样本程序合成,经闭合环路自修复与持久技能记忆,到极少数系统(如ASPIRE、ENPIRE、RoboClaw)实现的开放循环——执行反馈、技能记忆与演化搜索三者融合。同时梳理了技能端的互补方向,涵盖无监督强化学习技能发现至大语言模型技能库,并指出‘技能’至少有五种不同含义,唯有代码意义的技能能在不依赖梯度更新下实现自我改进。最后关联新兴技能经济:商业机器人技能市场已支持一键分发,但仅提供静态回放,暴露出适应性、跨具身迁移、溯源、安全验证、组合与标准化等关键挑战。本综述聚焦明确,通过一个分类体系与对比表,考察6类共77个代表性系统,给出自我改进机制的操作定义及各技术家族的局限性说明。

原文摘要 · Abstract (English)

Robot learning is splitting into two bets: policies that bake competence into frozen weights (vision-language-action, or VLA, models), and agents that write and refine their own executable skills as code. This survey organises the field around that axis of weights versus skills. Its central analytical contribution is a deep-dive that arranges code-as-policy methods by their degree of self-improvement, from zero-shot program synthesis, through closed-loop self-repair and persistent skill memory, to the sparsely populated cell in which execution feedback, skill memory, and evolutionary search combine into one open-ended loop; only a few very recent systems (for example ASPIRE, ENPIRE, and RoboClaw) occupy that cell. We map the complementary "skills" pole, from unsupervised reinforcement-learning skill discovery to large-language-model skill libraries, and show that the word "skill" is used in at least five distinct senses, of which only the code sense self-improves without gradient updates. We then connect the taxonomy to the emerging skill economy: commercial robot-skill marketplaces now distribute one-tap skills across robots but ship only static playback, which surfaces open problems of adaptation, cross-embodiment portability, provenance, safety verification, composition, and standardisation. This is a deliberately focused survey. Rather than cataloguing the field exhaustively, it examines 77 representative systems across six technique families through one taxonomy and a set of contrast tables, and it supplies operational definitions of the self-improvement mechanisms together with a statement of what each family cannot do.

机器人学习代码即策略技能经济自进化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。