arXiv:2602.18803cs.RO2026-02

通过图像空间定位参考轨迹,实现无需校准的跨机器人导航

Learning to Localize Reference Trajectories in Image-Space for Visual Navigation

  • 在当前视图中定位参考轨迹的图像坐标,不依赖机器人型号或标定
  • 在多种仿真与真实环境中成功率达94%-98%,比现有方法高20-50个百分点
  • 支持零样本迁移,手机视频即可让不同机器人沿轨迹导航

我们提出LoTIS,一种视觉导航模型,通过在机器人当前视图中定位参考RGB轨迹的图像空间坐标,提供与机器人无关的视觉引导,无需相机标定、位姿信息或特定机器人训练。不同于预测与具体机器人绑定的动作,该模型直接预测参考轨迹点在当前视角下的图像坐标,从而生成可通用的视觉指引,并轻松集成到局部规划中。通过将感知与动作解耦,并学习定位轨迹点而非模仿行为先验,实现了跨轨迹训练策略,增强对视角和相机变化的鲁棒性。在常规前向导航任务中,成功率提升20-50个百分点,达到94%-98%;在后向穿越等挑战性任务上,性能提升超过5倍。系统使用简单:我们展示仅用手机拍摄的视频,即可使不同机器人导航至轨迹任意点。视频、演示及代码已公开于https://finnbusch.com/lotis。

原文摘要 · Abstract (English)

We present LoTIS, a model for visual navigation that provides robot-agnostic image-space guidance by localizing a reference RGB trajectory in the robot's current view, without requiring camera calibration, poses, or robot-specific training. Instead of predicting actions tied to specific robots, we predict the image-space coordinates of the reference trajectory as they would appear in the robot's current view. This creates robot-agnostic visual guidance that easily integrates with local planning. Consequently, our model's predictions provide guidance zero-shot across diverse embodiments. By decoupling perception from action and learning to localize trajectory points rather than imitate behavioral priors, we enable a cross-trajectory training strategy for robustness to viewpoint and camera changes. We outperform state-of-the-art methods by 20-50 percentage points in success rate on conventional forward navigation, achieving 94-98% success rate across diverse sim and real environments. Furthermore, we achieve over 5x improvements on challenging tasks where baselines fail, such as backward traversal. The system is straightforward to use: we show how even a video from a phone camera directly enables different robots to navigate to any point on the trajectory. Videos, demo, and code are available at https://finnbusch.com/lotis.

视觉导航图像定位跨机器人零样本迁移

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。