让机器人导航更稳定,通过动作历史感知实现跨场景零样本部署。
VISTA: Scale-Aware Visual Navigation via Action History Conditioning

- 用动作历史条件化模型,显式建模预测与实际位移的关系。
- 在户外、森林、办公室等新环境中实现100%目标达成率。
- 适合需要高鲁棒性的真实机器人导航任务,尤其视觉重复场景。
视觉导航基础模型(VNMs)有望实现端到端学习的导航策略,支持跨不同机器人形态和环境的零样本部署。为保持通用性,多数基于视觉的导航模型预测归一化动作,但这种归一化引入了关键部署风险:对同一归一化轨迹应用不同缩放因子会改变其物理几何,导致性能下降和碰撞风险上升。为此,我们通过将模型对齐图像观测与归一化动作历史进行条件化,显式提供模型预测与机器人实际位移间关系的上下文信息。此外,现有VNMs在缺乏明显特征的视觉重复环境中表现不佳。为解决此问题,我们集成DINOv3编码器,其更丰富的表征使模型能捕捉观察间的空间与几何关系。VISTA在分布外环境中表现出强泛化能力,在室外、森林和办公场景的零样本真实部署中达到100%目标预测准确率,平均95%检查点成功通过,展现出在未见环境中的稳定路径遵循能力。
原文摘要 · Abstract (English)
Vision Navigation Foundation Models (VNMs) promise end-to-end learned navigation policies capable of zero-shot deployment across diverse embodiments and environments. To maintain generality, many vision-based navigation models predict normalized actions. However, this normalization introduces a critical deployment vulnerability: applying different scaling factors to the same normalized trajectory alters its physical geometry, which degrades navigation performance and increases collision risks. We address this vulnerability by conditioning the model on normalized action histories alongside image observations, providing explicit context on the relationship between the model's predictions and the robot's actual physical displacement. Furthermore, current VNMs often struggle in visually repetitive environments that lack distinct features. To resolve this issue, we integrate a DINOv3 encoder, whose richer representations enable our model to capture both spatial and geometric dimensions between observations. VISTA generalizes robustly to out-of-distribution environments, achieving 100% goal prediction accuracy in zero-shot, real-world deployment in Outdoor, Forest and Office settings, and an average of 95% checkpoints crossed, demonstrating consistent path following in unseen environments.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。