arXiv:2502.03729cs.ROcs.AI2025-02被引 27

用无动作人类视频训练机器人,通过语言推理实现跨形态泛化。

Action-Free Reasoning for Policy Generalization

  • 用人类视频的语义推理替代动作标签,学习可迁移的决策逻辑。
  • 在3,377条人类手部示范上训练,性能随推理数据量提升显著。
  • 适合研究机器人泛化、少标注学习与基于推理的控制方向。

端到端模仿学习为训练机器人策略提供了有前景的方法,但泛化到新环境仍是重大挑战。尽管大规模机器人示范数据展现出泛化潜力,但其扩展成本高昂。相比之下,人类视频数据丰富多样,是极具吸引力的替代方案。然而,这些数据缺乏动作标签,难以直接用于模仿学习。现有方法尝试从视频中提取具身动作表示(如手部姿态),但生成的策略仍难以弥合人与机器人动作间的具身差距。本文提出一种新思路:利用人类视频中的语言推理(对机器人动作至关重要的指导信息)来训练可泛化的机器人策略。基于近期基于推理的策略架构进展,我们提出行动无关数据推理(RAD)。RAD同时学习机器人示范数据(含推理与动作标签)和无动作的人类视频数据(仅含推理标签)。前者教会模型将推理映射为底层动作,后者增强推理能力。此外,我们将发布一个包含3,377条人类手部示范的新数据集,附带推理标注,兼容Bridge V2基准,以推动未来基于推理的机器人学习研究。实验表明,RAD能有效跨越具身差距,使机器人执行仅在无动作数据中见过的任务。进一步地,扩大无动作推理数据规模可显著提升策略性能与对新任务的泛化能力。这些结果凸显了从无动作数据中进行推理驱动学习在推进可泛化机器人控制方面的潜力。

原文摘要 · Abstract (English)

End-to-end imitation learning offers a promising approach for training robot policies. However, generalizing to new settings remains a significant challenge. Although large-scale robot demonstration datasets have shown potential for inducing generalization, they are resource-intensive to scale. In contrast, human video data is abundant and diverse, presenting an attractive alternative. Yet, these human-video datasets lack action labels, complicating their use in imitation learning. Existing methods attempt to extract grounded action representations (e.g., hand poses), but resulting policies struggle to bridge the embodiment gap between human and robot actions. We propose an alternative approach: leveraging language-based reasoning from human videos-essential for guiding robot actions-to train generalizable robot policies. Building on recent advances in reasoning-based policy architectures, we introduce Reasoning through Action-free Data (RAD). RAD learns from both robot demonstration data (with reasoning and action labels) and action-free human video data (with only reasoning labels). The robot data teaches the model to map reasoning to low-level actions, while the action-free data enhances reasoning capabilities. Additionally, we will release a new dataset of 3,377 human-hand demonstrations with reasoning annotations compatible with the Bridge V2 benchmark and aimed at facilitating future research on reasoning-driven robot learning. Our experiments show that RAD enables effective transfer across the embodiment gap, allowing robots to perform tasks seen only in action-free data. Furthermore, scaling up action-free reasoning data significantly improves policy performance and generalization to novel tasks. These results highlight the promise of reasoning-driven learning from action-free datasets for advancing generalizable robot control. Project page: https://rad-generalization.github.io

机器人控制推理学习无动作数据泛化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。