用时空事件图提升多模态长期导航的可靠性和任务复用能力
STEGNav: Spatio-Temporal Event Graph Reasoning for Multimodal Lifelong Object Navigation

- 构建时空事件图,融合空间实例定位与时间轨迹记忆
- 在GOAT-Bench上达66.3%成功率和39.7%SPL
- 无需训练,适合长期探索与跨任务学习场景
多模态长期导航要求智能体在未见过环境中自主探索,并按物体类别、语言描述或参考图像顺序完成导航任务。现有方法主要依赖以状态为中心的语义场景图,但难以区分相似实例,无法联合表示语义目标与探索前沿,且难有效利用导航记忆与轨迹经验。为此,我们提出无训练的时空事件图导航(STEGNav)框架,将传统场景图扩展为沿空间与时间轴的互补时空事件图。空间轴通过查询条件实例定位,联合表征语义目标与可达性、路径代价、探索效用驱动的占用感知探索前沿;时间轴采用轨迹感知双窗口记忆,保留近期决策-轨迹事件及跨子任务验证结果。基于视觉语言模型的导航代理在生成的时空事件图上推理,选择目标实例或探索前沿作为下一步目标。STEGNav在GOAT-Bench上实现66.3%成功率和39.7%SPL,HM3Dv1和HM3Dv2上分别达到64.0%和69.4%成功率。消融实验与错误分析验证了两轴互补作用,表明事件驱动的时空表征提升了导航可靠性与跨子任务经验复用能力。
原文摘要 · Abstract (English)
Multimodal lifelong navigation requires an agent to autonomously explore unseen environments while sequentially completing navigation tasks specified by object categories, language descriptions, or reference images. Existing methods primarily accomplish these tasks by constructing state-centric semantic scene graphs. By treating scene graphs as persistent repositories of semantic observations, these methods struggle to distinguish similar instances, jointly represent semantic targets and exploration frontiers, and effectively exploit navigation memory and trajectory experience. To address these limitations, we propose Spatio-Temporal Event Graph Navigation (STEGNav), a training-free framework that extends conventional scene graphs into spatio-temporal event graphs along complementary spatial and temporal axes. The spatial axis performs query-conditioned instance grounding and jointly represents semantic targets and occupancy-aware exploration frontiers characterized by reachability, path cost, and exploration utility. The temporal axis employs trajectory-aware dual-window memory to retain recent decision--trajectory events and verified cross-subtask navigation outcomes. A VLM-based navigation agent reasons over the resulting spatio-temporal event graph and selects either a target instance or an exploration frontier as its next navigation goal. STEGNav achieves 66.3% SR and 39.7 SPL on GOAT-Bench, as well as SR scores of 64.0% and 69.4% on HM3Dv1 and HM3Dv2, respectively. Ablation studies and error analyses validate the complementary effects of the two axes, demonstrating that event-driven spatio-temporal representations improve navigation reliability and cross-subtask experience reuse.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。