arXiv:2512.20052cs.AIcs.RO2025-12被引 2

从无动作视频中学习机器人技能,实现高层规划与动作转换。

Learning Skills from Action-Free Videos

  • 基于光流构建运动表征,学习视频中的潜在技能空间。
  • 在多任务和长序列任务中性能显著提升,可直接从视觉数据获取技能。
  • 适合希望无需标注动作即训练通用机器人的研究者。

视频学习为通用机器人提供了有前景的路径,因其包含超越真实机器人数据集的丰富视觉与时间先验。现有视频生成模型虽能生成逼真视觉预测,但难以转化为低层动作;而潜在动作模型通常仅支持单步操作,缺乏高层规划能力。本文提出从光流中学习技能(SOF)框架,通过光流构建的中间表示,学习大规模无动作视频中的潜在技能空间,该表示同时捕捉视频动态与机器人动作的运动信息。在流形空间中学习技能,使高层规划成为可能,并简化技能到动作的映射。实验表明,该方法在多任务与长程任务中均表现更优,证明了可直接从原始视觉数据中学习并组合技能。

原文摘要 · Abstract (English)

Learning from videos offers a promising path toward generalist robots by providing rich visual and temporal priors beyond what real robot datasets contain. While existing video generative models produce impressive visual predictions, they are difficult to translate into low-level actions. Conversely, latent-action models better align videos with actions, but they typically operate at the single-step level and lack high-level planning capabilities. We bridge this gap by introducing Skill Abstraction from Optical Flow (SOF), a framework that learns latent skills from large collections of action-free videos. Our key idea is to learn a latent skill space through an intermediate representation based on optical flow that captures motion information aligned with both video dynamics and robot actions. By learning skills in this flow-based latent space, SOF enables high-level planning over video-derived skills and allows for easier translation of these skills into actions. Experiments show that our approach consistently improves performance in both multitask and long-horizon settings, demonstrating the ability to acquire and compose skills directly from raw visual data.

视频学习技能学习机器人

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。