arXiv:2502.09886cs.ROcs.AI2025-02被引 29

用网络视频生成机器人任务,低成本训练通用操作策略。

Video2Policy: Scaling up Manipulation Tasks in Simulation through Internet Videos

  • 从真实人类视频重建仿真任务,无需人工标注。
  • 在9类任务上成功训练出可执行投掷等复杂动作的策略。
  • 适合需要大量多样化训练数据的机器人研究者。

仿真为通用策略的低成本数据扩展提供了前景。为可扩展地生成多样且真实的任务数据,现有方法或依赖可能虚构不相关任务的大语言模型(LLMs),或依赖需精细真实-仿真对齐的数字孪生,难以规模化。为此,我们提出Video2Policy,一种利用互联网RGB视频重构日常人类行为任务的新框架。该方法包含两阶段:(1) 从视频生成仿真中的任务;(2) 利用上下文内LLM生成的奖励函数进行迭代强化学习。我们在Something-Something-v2(SSv2)数据集上重建超过100段视频,涵盖9种不同任务,展示其多样性与复杂性。该方法成功训练出可在这些任务上运行的强化学习策略,包括投掷等复杂动作。最后,我们验证生成的仿真数据可规模化用于训练通用策略,并可通过Real2Sim2Real方式迁移回真实机器人。

原文摘要 · Abstract (English)

Simulation offers a promising approach for cheaply scaling training data for generalist policies. To scalably generate data from diverse and realistic tasks, existing algorithms either rely on large language models (LLMs) that may hallucinate tasks not interesting for robotics; or digital twins, which require careful real-to-sim alignment and are hard to scale. To address these challenges, we introduce Video2Policy, a novel framework that leverages internet RGB videos to reconstruct tasks based on everyday human behavior. Our approach comprises two phases: (1) task generation in simulation from videos; and (2) reinforcement learning utilizing in-context LLM-generated reward functions iteratively. We demonstrate the efficacy of Video2Policy by reconstructing over 100 videos from the Something-Something-v2 (SSv2) dataset, which depicts diverse and complex human behaviors on 9 different tasks. Our method can successfully train RL policies on such tasks, including complex and challenging tasks such as throwing. Finally, we show that the generated simulation data can be scaled up for training a general policy, and it can be transferred back to the real robot in a Real2Sim2Real way.

机器人学习仿真训练视频理解强化学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。