arXiv:2507.13340cs.ROcs.AI2025-07被引 9

用视觉流统一不同机器人动作,低数据下提升策略性能。

Latent Policy Steering with Embodiment-Agnostic Pretrained World Models

  • 用光流作为通用动作表示预训练世界模型,跨形态数据可复用。
  • 真实世界实验中仅需30-50次演示即达70%性能提升。
  • 适合缺乏大量目标机器人数据的现实部署场景。

学习型机器人视觉运动策略的表现高度依赖训练数据规模与质量。尽管大规模机器人和人类数据集日益可用,但实体差异和动作空间不匹配使其难以利用。本文核心洞察是:不同实体执行的技能在动作视觉上具有相似性,可通过现成的动作表示(如光流)捕捉。此外,世界模型(WM)能利用次优数据,因其专注于建模动态。本工作通过先使用光流作为无体感动作表示,在多源实体(机器人、人类)的易获取数据上预训练一个世界模型。随后在少量目标实体示范数据上微调该模型,对齐预测结果,训练基础策略并学习鲁棒价值函数。基于微调后的世界模型与价值函数,方法评估基础策略生成的动作候选,选择最优动作以提升性能。所提方法称为潜在策略引导(LPS),在四个Robomimic任务上平均提升行为克隆策略10.6%,且多数预训练数据来自真实世界。真实实验中,仅需30-50次目标演示即实现70%相对提升,60-100次演示时达44%相对提升,显著优于行为克隆基线。

原文摘要 · Abstract (English)

The performance of learned robot visuomotor policies is heavily dependent on the size and quality of the training dataset. Although large-scale robot and human datasets are increasingly available, embodiment gaps and mismatched action spaces make them difficult to leverage. Our main insight is that skills performed across different embodiments produce visual similarities in motions that can be captured using off-the-shelf action representations such as optical flow. Moreover, World Models (WMs) can leverage sub-optimal data since they focus on modeling dynamics. In this work, we aim to improve visuomotor policies in low-data regimes by first pretraining a WM using optical flow as an embodiment-agnostic action representation to leverage accessible or easily collected data from multiple embodiments (robots, humans). Given a small set of demonstrations on a target embodiment, we finetune the WM on this data to better align the WM predictions, train a base policy, and learn a robust value function. Using our finetuned WM and value function, our approach evaluates action candidates from the base policy and selects the best one to improve performance. Our approach, which we term Latent Policy Steering (LPS), improves behavior-cloned policies by 10.6% on average across four Robomimic tasks, even though most of the pretraining data comes from the real world. In the real-world experiments, LPS achieves larger gains: 70% relative improvement with 30-50 target-embodiment demonstrations, and 44% relative improvement with 60-100 demonstrations, compared to a behavior-cloned baseline. Qualitative results can be found on the website: https://yiqiwang8177.github.io/LatentPolicySteering/.

视觉运动世界模型低数据跨形态

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。