arXiv:2606.01951cs.RO2026-06

用人类第一视角视频训练机器人导航,提升语言理解与动作生成能力。

Co-training with Ego-centric Video and Demonstration for Robot Navigation Task

论文配图:Co-training with Ego-centric Video and Demonstration for Robot Navigation Task
图 1 · 摘自论文原文
  • 从人类行走视频中估计相机运动,转为机器人可执行的动作表示
  • 联合训练使模型在水果搜索任务上表现优于单一数据源
  • 适合需要低成本数据扩展的移动机器人学习场景

视觉-语言-动作(VLA)模型在多样化机器人任务中前景广阔,但其性能高度依赖大规模高质量训练数据,而真实机器人采集此类数据成本高、耗时长。以往研究曾尝试用第一视角人类视频扩充操作类数据集,但将该方法应用于移动机器人导航仍面临运动过程中视角变化的挑战。本文提出一种框架,将人类第一视角行走视频转换为适用于移动机器人模仿学习的数据集。该方法通过分析人类视频中的相机运动,将其转化为与地面机器人兼容的动作表示。通过在人类衍生数据与机器人实测数据上联合训练VLA模型,模型在语言理解与动作生成方面均优于仅使用单一数据源的情况。在水果搜索导航任务上的实验表明,人类第一视角视频为移动机器人学习提供了有效且可扩展的数据来源。

原文摘要 · Abstract (English)

Vision-language-action (VLA) models are promising for diverse robotic tasks, but their performance heavily depends on large-scale high-quality training data, whose collection on real robots is costly and time-consuming. While prior work has explored augmenting manipulation datasets with egocentric human videos, applying such approaches to mobile robot navigation remains challenging due to viewpoint changes during locomotion. In this paper, we propose a framework that converts egocentric walking videos into datasets for mobile robot imitation learning. The proposed method estimates camera motion from human videos and transforms it into action representations compatible with ground mobile robots. By jointly training a VLA model on human-derived and robot-collected datasets, the model achieves improved language understanding and more robust action generation than training with either data source alone. Experiments on a fruit-search navigation task demonstrate that human egocentric videos provide an effective and scalable data source for mobile robot learning.

机器人导航第一视角视频模仿学习VLA模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。