arXiv:2602.10594cs.ROcs.LG2026-02中稿 · ICRA

用人体视频学机器人动作,少样本也能精准模仿。

Flow-Enabled Generalization to Human Demonstrations in Few-Shot Imitation Learning

  • 通过场景光流预测跨体感动作轨迹,融合人机视频数据。
  • 在真实任务中优于现有方法,仅需少量演示即实现高精度控制。
  • 适合需要少样本、跨场景泛化的机器人模仿学习任务。

模仿学习(IL)使机器人无需显式建模即可从示范中学习复杂技能,但通常需大量示范,收集成本高。先前工作尝试使用光流作为中间表示,用人体视频替代部分机器人示范,以减少采集量。然而,多数方法仅关注物体或机械臂/手特定点的光流,无法描述交互运动;且依赖光流实现对仅在人体视频中出现场景的泛化能力有限,因光流难以捕捉精确运动细节。此外,基于场景观测生成动作可能导致光流条件策略过拟合训练任务,削弱光流带来的泛化能力。为此,我们提出 SFCrP,包含用于跨体感学习的场景光流预测模型(SFCr)和光流与裁剪点云条件策略(FCrP)。SFCr 从机器人与人体视频中学习,预测任意点轨迹;FCrP 遵循整体光流运动趋势,并根据观测调整动作以实现精度控制。所提方法在多种真实世界任务设置下超越当前最优基线,同时展现出对仅在人体视频中出现场景的强空间与实例泛化能力。

原文摘要 · Abstract (English)

Imitation Learning (IL) enables robots to learn complex skills from demonstrations without explicit task modeling, but it typically requires large amounts of demonstrations, creating significant collection costs. Prior work has investigated using flow as an intermediate representation to enable the use of human videos as a substitute, thereby reducing the amount of required robot demonstrations. However, most prior work has focused on the flow, either on the object or on specific points of the robot/hand, which cannot describe the motion of interaction. Meanwhile, relying on flow to achieve generalization to scenarios observed only in human videos remains limited, as flow alone cannot capture precise motion details. Furthermore, conditioning on scene observation to produce precise actions may cause the flow-conditioned policy to overfit to training tasks and weaken the generalization indicated by the flow. To address these gaps, we propose SFCrP, which includes a Scene Flow prediction model for Cross-embodiment learning (SFCr) and a Flow and Cropped point cloud conditioned Policy (FCrP). SFCr learns from both robot and human videos and predicts any point trajectories. FCrP follows the general flow motion and adjusts the action based on observations for precision tasks. Our method outperforms SOTA baselines across various real-world task settings, while also exhibiting strong spatial and instance generalization to scenarios seen only in human videos.

模仿学习光流少样本跨体感

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。