构建可跨楼层持续记忆的3D导航系统,实现多目标顺序寻物。
LifelongCrossNav: Persistent 3D Semantic Memory for Cross-Floor Multi-Object Navigation

- 用稀疏3D语义体素存储环境结构与视觉语言特征
- 跨楼层导航准确率提升,无需重建地图
- 适合需要长期记忆和多层探索的智能机器人
物体目标导航在语义感知和探索方面已取得显著进展,但多物体导航与跨楼层导航仍常被单独处理。我们提出LifelongCrossNav,一种用于未知多层室内环境中连续多目标物体导航的框架。每轮任务中,智能体接收有序的物体目标查询序列,并持续维护一个共享的稀疏3D语义体素记忆。该记忆逐步累积几何结构、可通行状态及视觉-语言特征,使后续目标查询可复用已有场景信息而无需重新建图。为支持跨楼层搜索,LifelongCrossNav结合了感知自适应的3D可通行性映射、楼梯专用感知与方向感知的楼梯穿越策略。统一的导航策略协调同一层的前沿探索、实时与历史兴趣点检索、楼梯导航以及目标物体搜寻与接近。我们进一步构建了基于HM3D场景的基准数据集HM3D-MFMON,包含需至少一次楼层切换才能完成全部任务子任务的专属子集。实验表明,LifelongCrossNav在HM3D-MFMON上始终优于代表性平面持久语义地图基线,证明持久3D语义记忆与跨楼层可通行性建模能有效支撑多层环境中的连续多目标导航。
原文摘要 · Abstract (English)
Object-goal navigation has made substantial progress in semantic perception and exploration, yet persistent memory for multi-object navigation and cross-floor navigation are still commonly addressed separately. We present LifelongCrossNav, a framework for sequential multi-object ObjectNav in unknown multi-floor indoor environments. Within each episode, the agent receives an ordered sequence of object-goal queries while continuously maintaining a shared sparse 3D semantic voxel memory. This memory incrementally accumulates geometric structure, traversability states, and vision-language features, allowing subsequent object-goal queries to retrieve previously acquired scene information without rebuilding the map. To support persistent search across floors, LifelongCrossNav combines support-aware 3D traversability mapping, stair-specific perception, and direction-aware stair traversal. A unified navigation policy coordinates same-floor frontier exploration, live and historical point-of-interest retrieval, stair navigation, and target-object search and approach. We further introduce HM3D-MFMON, a benchmark for sequential Multi-Floor Multi-Object Navigation built on HM3D scenes, including a dedicated subset in which completing the full sequence of object-goal subtasks requires at least one floor transition. Experimental results show that LifelongCrossNav consistently outperforms a representative planar persistent semantic-map baseline on HM3D-MFMON, demonstrating that persistent 3D semantic memory and cross-floor traversability modeling effectively support sequential multi-object navigation in multi-floor environments. Project page: https://flageval-baai.github.io/LifelongCrossNavPage.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。