arXiv:2606.01788cs.CV2026-06

用纯视觉构建通用语义地图,实现跨任务导航零训练

PlatonicNav: Unveiling Semantic Correspondence in Navigation with Platonic Topological Maps

论文配图:PlatonicNav: Unveiling Semantic Correspondence in Navigation with Platonic Topological Maps
图 1 · 摘自论文原文
  • 基于自监督视觉编码器构建融合几何与语义的拓扑地图
  • 无需配对图文数据,在多个任务上实现0训练泛化
  • 适合需要跨模态兼容性的机器人导航场景

具身视觉导航要求智能体从原始感官输入中感知复杂环境并执行动作抵达目标,广泛应用于家庭服务机器人、辅助机器人和大规模自主探索。然而,当前统一视觉-语言导航(VLN)与物体目标导航(ObjNav)的方法仍停留在架构融合、混合训练和大规模视觉-语言预训练层面,未探究独立训练的视觉与语言编码器是否已具备共同语义结构。此外,即便采用以物体为中心的拓扑地图,也依赖CLIP等模型进行显式跨模态监督来关联语言目标,尚未验证是否可仅通过纯视觉构建的地图完成语言目标定位。为此,我们首次将柏拉图表征假说扩展至具身导航领域,将视觉单模态ObjNav、跨模态ObjNav及VLN重新定义为同一物体中心语义流形的三种不同接口。我们提出PlatonicNav——一种无需训练的框架,其柏拉图拓扑地图通过自监督视觉编码器融合几何与语义节点距离,并通过盲匹配方式实现语言目标定位,不依赖任何成对的视觉-语言数据。在HM3D-IIN、OVON和R2R-CE(MP3D)等多个仿真基准上的大量实验,以及在Unitree Go2机器人的部署结果表明,PlatonicNav在任务、模态和机器人本体之间均表现出强泛化能力,且无需显式跨模态训练。

原文摘要 · Abstract (English)

Embodied visual navigation, where an agent perceives a complex environment and acts to reach a goal from raw sensory input, underpins a wide range of applications such as household service robotics, assistive robotics, and large-scale autonomous exploration. However, recent attempts to unify vision-and-language navigation (VLN) and object goal navigation (ObjNav) remain at the level of architectural fusion, mixed-task training, and large vision-language pretraining, without examining whether independently trained vision and language encoders may already share a common semantic structure. Moreover, even object-centric topological maps still ground language goals through explicit cross-modal supervision such as CLIP or large vision-language models, leaving open whether such grounding is possible from a purely vision-built map. To address these challenges, we extend the Platonic Representation Hypothesis to embodied navigation and recast vision-only ObjNav, cross-modal ObjNav, and VLN as three different interfaces to the same object-centric semantic manifold. We further introduce PlatonicNav, a training-free framework whose Platonic Topological Map fuses geometric and semantic node distances from a self-supervised visual encoder, and grounds language goals via blind matching without any paired vision-language data. Extensive experiments on simulation benchmarks including HM3D-IIN, OVON, and R2R-CE on MP3D, together with deployment on Unitree Go2, demonstrate that PlatonicNav generalizes across tasks, modalities, and embodiments without explicit cross-modal training. Code: https://github.com/AIGeeksGroup/PlatonicNav. Website: https://aigeeksgroup.github.io/PlatonicNav.

具身导航语义地图自监督零样本

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。