提升视觉语言导航的拓扑规划,聚焦关键路径点
LCGNav: Local Candidate-Aware Geometric Enhancement for General Topological Planning in Vision-Language Navigation

- 将候选视野转为3D点云,按可达范围裁剪,压缩局部几何信息
- 仅对当前相关节点增强几何特征,不改变原规划接口
- 在R2R-CE和RxR-CE上显著提升多个基线模型性能
在线拓扑规划已成为连续环境视觉语言导航(VLN-CE)的有效范式,但现有方法仍存在两个局限:冗余的局部深度信息,以及随着拓扑图增长,对当前前缘候选点的关注度减弱。为此,我们提出LCGNav,一种模块化的局部几何增强框架,用于拓扑视觉语言导航。LCGNav显式将候选深度视图转换为3D点云,并基于智能体可达范围进行物理截断,实现更紧凑的局部几何建模。同时引入保持维度的局部融合策略与瞬态状态退化机制,使几何增强仅作用于当前相关的虚节点,而不改变原始规划器接口。在R2R-CE和RxR-CE上的实验表明,LCGNav作为跨架构增强模块表现优异,以极低额外训练成本持续提升多个代表性在线拓扑基线的关键指标。集成ETP-R1后,其在R2R-CE和RxR-CE验证未见分割上达到当前最优性能。代码已公开于https://github.com/shannanshouyin/LCGNav。
原文摘要 · Abstract (English)
Online topological planning has become an effective paradigm for Vision-Language Navigation in Continuous Environments (VLN-CE), but existing methods still suffer from two limitations: redundant local depth information and weakened focus on current frontier candidates as the topological graph grows. To address this, we propose LCGNav, a modular local geometric enhancement framework for topological VLN. LCGNav explicitly converts candidate depth views into 3D point clouds and applies physical truncation based on the agent's reachable range, enabling more compact local geometric modeling. It further introduces a dimension-preserving local fusion strategy with transient state degradation, so that geometric enhancement is applied only to the currently relevant ghost nodes without changing the original planner interface. Experiments on R2R-CE and RxR-CE show that LCGNav serves as an effective cross-architecture enhancement module, consistently improving multiple key metrics of representative online topological baselines with low additional training cost. When integrated with ETP-R1, LCGNav achieves the best performance among the compared online topological methods on the val-unseen splits of the R2R-CE and RxR-CE benchmarks. The code is available at https://github.com/shannanshouyin/LCGNav.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。