提出新基准评测智能体在变化环境中的长期导航能力
EvoNav-Bench: Benchmarking Lifelong Navigation in Evolving Environments

- 构建动态环境下的长期导航评测框架,模拟真实世界场景演化
- 现有方法在环境变化下表现明显下降,依赖旧经验易出错
- 验证简单启发式策略可有效应对环境变迁,适合研究鲁棒性
长期导航(LN)要求智能体在相同环境中连续完成多个导航任务。由于每次从头开始探索会带来冗余,LN智能体必须积累并复用早期经验,通常通过持久化场景表示如场景图或视觉快照实现。然而,现有方法多假设环境静态,而现实中的长期导航中,人类活动会导致环境持续变化。当这一假设失效时,现有方法可能将过时的先验观察与新观测融合,但当前基准无法揭示此失败模式。本文提出EvoNav-Bench,基于ProcTHOR框架,在导航任务间引入环境修改,使历史经验仍有用但不再完全可靠。该设计支持对环境演化如何影响依赖过往观测的LN智能体进行可控评估。我们使用EvoNav-Bench测试了三种近期构建并复用场景表示的方法,并对比了三种简单启发式策略:前沿更新、失败后更新、阶段重置。结果表明,现有方法在环境演化下脆弱,而启发式策略能有效分析智能体适应变化的能力并减轻其影响。
原文摘要 · Abstract (English)
Lifelong navigation (LN) requires an embodied agent to solve a sequence of navigation subtasks in the same environment. Since solving each subtask from scratch incurs redundant exploration, an LN agent must consolidate experience from earlier stages and reuse it in later stages, often through persistent scene representations such as scene graphs or visual snapshots. However, existing approaches typically assume a stationary environment, whereas in real-world LN settings, human activities can cause the environment to evolve. With the stationary assumption violated, existing methods may fuse outdated prior observations with new observations, yet current benchmarks cannot reveal this failure mode. In this paper, we present EvoNav-Bench, which extends the GOAT-Bench style LN formulation in the context of evolving environments. Built on the ProcTHOR framework, EvoNav-Bench introduces environment modifications between navigation tasks, making prior experience useful but not fully reliable. This design enables controlled evaluation of how environment evolution affects LN agents that reuse prior scene observations. Using EvoNav-Bench, we benchmark three recent methods that build and reuse scene representations for navigation. We also compare three simple heuristic strategies for handling environment evolution: Frontier-Update, Fail-then-Update, and Stage-Reset. Our results show that existing methods are brittle under environment evolution, while the heuristic strategies enable a controlled analysis of how agents can adapt to scene changes and mitigate their impact.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。