用高层语义潜空间实现高效图像目标导航,省去冗余像素重建。
Efficient Image-Goal Navigation with Representative Latent World Model
- 在预训练的语义潜空间中直接规划,避免像素级重建开销。
- 在多个基准上达到顶尖轨迹预测与图像目标导航性能。
- 已在真实机器人上验证,适合对效率要求高的实际应用。
世界模型通过预测未来世界状态,使机器人能在物理环境中进行反事实推理。传统方法通常注重未来场景的像素级重建,但这种高精度渲染计算成本高昂,且对导航等规划任务并非必需。为此,我们提出在高层语义表示的潜空间中直接进行预测与规划。为此,我们引入代表性潜空间导航世界模型(ReL-NWM)。该方法不依赖重建导向的潜空间编码,而是利用预训练的表示编码器DINOv3,并结合专门机制,在表示空间中有效融合动作信号与历史上下文信息。整个模型完全在潜空间内运行,跳过昂贵的显式重建步骤,实现高效的导航规划。实验表明,该模型在多个基准上均达到最先进的轨迹预测与图像目标导航性能。此外,我们在Unitree G1人形机器人上部署系统,验证了其在实际导航场景中的高效性与鲁棒性。
原文摘要 · Abstract (English)
World models enable robots to conduct counterfactual reasoning in physical environments by predicting future world states. While conventional approaches often prioritize pixel-level reconstruction of future scenes, such detailed rendering is computationally intensive and unnecessary for planning tasks like navigation. We therefore propose that prediction and planning can be efficiently performed directly within a latent space of high-level semantic representations. To realize this, we introduce the Representative Latent space Navigation World Model (ReL-NWM). Rather than relying on reconstructionoriented latent embeddings, our method leverages a pre-trained representation encoder, DINOv3, and incorporates specialized mechanisms to effectively integrate action signals and historical context within this representation space. By operating entirely in the latent domain, our model bypasses expensive explicit reconstruction and achieves highly efficient navigation planning. Experiments show state-of-the-art trajectory prediction and image-goal navigation performance on multiple benchmarks. Additionally, we demonstrate real-world applicability by deploying the system on a Unitree G1 humanoid robot, confirming its efficiency and robustness in practical navigation scenarios.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。