预测真人视角下人形机器人头部6自由度姿态,实现真实场景自主导航。
LookOut: Real-World Humanoid Egocentric Navigation
- 基于时序聚合的3D潜在特征建模环境几何与语义约束
- 在4小时真实场景数据上训练,可泛化至未见环境
- 学习人类导航行为如环顾、减速、绕行,适合机器人视觉研究
从第一人称视频中预测无碰撞的未来6自由度头部姿态,在人形机器人、虚拟/增强现实及辅助导航中至关重要。本文提出从第一人称视频序列中预测头部平移与旋转的任务,以捕捉通过头部转动表达的主动信息获取行为。为此,我们设计一个基于时序聚合3D潜在特征的框架,建模静态与动态环境的几何和语义约束。针对该领域训练数据匮乏问题,我们利用Project Aria眼镜构建数据采集流程,并发布名为Aria Navigation Dataset(AND)的数据集,包含4小时真实场景下用户导航的记录,涵盖多样化情境与行为。大量实验表明,模型可学习人类式导航行为,如等待、减速、重规划路径以及观察交通状况,并在未见环境中良好泛化。
原文摘要 · Abstract (English)
The ability to predict collision-free future trajectories from egocentric observations is crucial in applications such as humanoid robotics, VR / AR, and assistive navigation. In this work, we introduce the challenging problem of predicting a sequence of future 6D head poses from an egocentric video. In particular, we predict both head translations and rotations to learn the active information-gathering behavior expressed through head-turning events. To solve this task, we propose a framework that reasons over temporally aggregated 3D latent features, which models the geometric and semantic constraints for both the static and dynamic parts of the environment. Motivated by the lack of training data in this space, we further contribute a data collection pipeline using the Project Aria glasses, and present a dataset collected through this approach. Our dataset, dubbed Aria Navigation Dataset (AND), consists of 4 hours of recording of users navigating in real-world scenarios. It includes diverse situations and navigation behaviors, providing a valuable resource for learning real-world egocentric navigation policies. Extensive experiments show that our model learns human-like navigation behaviors such as waiting / slowing down, rerouting, and looking around for traffic while generalizing to unseen environments. Check out our project webpage at https://sites.google.com/stanford.edu/lookout.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。