arXiv:2607.02466cs.ROcs.AI2026-07中稿 · ICML

用无监督交互数据先学动作技能,再少量标注就实现强泛化。

Learning to Move Before Learning to Do: Task-Agnostic pretraining for VLAs

论文配图:Learning to Move Before Learning to Do: Task-Agnostic pretraining for VLAs
图 1 · 摘自论文原文
  • 先用自监督学通用运动规律,再用少量专家数据对齐语言。
  • 仅用极少标注数据,在SIMPLER上达到百万级数据训练模型的效果。
  • 真实机器人实验中抗摄像头扰动能力提升25%,适合做具身智能预训练。

视觉-语言-动作(VLA)模型受限于专家示范数据稀缺——这些包含观测、指令和动作的三元组难以大规模获取。我们提出,这一瓶颈源于混淆了两类学习目标:掌握物理操作能力(如何移动)与语义对齐能力(做什么)。关键在于,后者才需要语言监督。基于分解假说,我们提出任务无关预训练(TAP),分两阶段进行:第一阶段通过自监督逆动力学目标,从廉价的未标注交互数据(包括被丢弃的非任务轨迹和自主机器人游戏)中学习可迁移的运动先验;第二阶段仅需少量专家数据,将这些先验与语言对齐。在SIMPLER基准上,TAP仅用极少量标注数据即达到超过100万条专家轨迹训练模型的性能,相比标准行为克隆提升10%。在真实世界WidowX平台上,当摄像头扰动下,互联网规模基线模型成功率为0%,而TAP仍保持25%成功率,表明任务无关预训练能生成鲁棒、可迁移的物理表征,为具身智能提供可扩展路径。

原文摘要 · Abstract (English)

Vision-Language-Action (VLA) models are fundamentally bottlenecked by the scarcity of expert demonstrations -- triplets of observations, instructions, and actions that are costly to collect at scale. We argue that this bottleneck stems from conflating two distinct learning objectives: acquiring physical competence (how to move) and acquiring semantic alignment (what to do). Crucially, only the latter requires language supervision. Building on this Decomposition Hypothesis, we propose Task-Agnostic Pretraining (TAP), a two-stage framework that first learns transferable motor priors from cheap, unlabeled interaction data -- including discarded off-task trajectories and autonomous robot play -- via a self-supervised Inverse Dynamics objective. A lightweight second stage then grounds these priors in language using minimal expert data. On the SIMPLER benchmark, TAP matches models trained on over 1M expert trajectories while using orders of magnitude less labeled data, yielding a 10% absolute gain over standard behavior cloning. On a real-world WidowX platform, TAP retains 25% success under camera perturbations where internet-scale baselines collapse to 0%, demonstrating that task-agnostic pretraining produces robust, transferable physical representations and offers a scalable path forward for Embodied AI.

具身智能预训练自监督机器人

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。