arXiv:2603.00167cs.RO2026-03被引 2

用短时视角视频预测全局动态地图,让机器人提前规划路径。

EgoMoD: Predicting Global Maps of Dynamics from Local Egocentric Observations

  • 基于局部视觉与姿态信息,端到端预测全局运动趋势。
  • 在有限观测下仍能准确预测未来动态分布,误差低于基准方法23%。
  • 无需外部全局传感,可直接部署于普通机器人上。

在动态环境中高效导航需预判机器人感知范围之外的运动演化,实现前瞻而非仅反应式规划。动态地图(MoDs)为长期全局规划提供了空间运动倾向的结构化表示,但传统构建方式依赖长时间的全局环境观测。本文提出EgoMoD,首个从机器人运行中采集的短时全景视频片段直接预测未来MoDs的方法。该方法通过视频与姿态条件化的架构,利用外部观测计算出的MoDs作为特权监督信号,学习从局部动态线索推断全局运动趋势的能力,使局部观测成为全局运动结构的预测信号。由此,EgoMoD可在不依赖过去模式外推的前提下,对整个环境的未来运动动态进行预测。作为场景特异的动态先验,它在推理阶段替代了以往方法所需的外部全局感知基础设施,仅使用标准机载传感器即可实现。大规模模拟环境实验表明,即使在观测受限条件下,EgoMoD仍能有效预测未来MoDs;真实图像评估进一步验证其零样本迁移至真实系统的能力。

原文摘要 · Abstract (English)

Efficient navigation in dynamic environments requires anticipating how motion patterns evolve beyond the robot's immediate perceptual range, enabling preemptive rather than purely reactive planning in crowded scenes. Maps of Dynamics (MoDs) offer a structured representation of motion tendencies in space useful for long-term global planning, but constructing them traditionally requires global environment observations over extended periods of time. We introduce EgoMoD, the first approach that learns to predict future MoDs directly from short egocentric video clips collected during robot operation. Our method learns to infer environment-wide motion tendencies from local dynamic cues using a video- and pose-conditioned architecture trained with MoDs computed from external observations as privileged supervision, allowing local observations to serve as predictive signals of global motion structure. Thanks to this, we offer the capacity to forecast future motion dynamics over the whole environment rather than merely extend past patterns in the robot's field of view. As a site-specific dynamic prior, EgoMoD replaces the external global sensing infrastructure required by prior MoD methods at inference time with standard onboard sensors. Experiments in large simulated environments show that EgoMoD predicts future MoDs under limited observability, while evaluation with real images showcases its zero-shot transferability to real systems.

动态地图机器人导航视觉预测端到端

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。