MUVLA通过地图理解实现物体导航的高效探索。
MUVLA: Learning to Explore Object Navigation via Map Understanding
- 利用语义地图统一历史信息,结构化空间上下文。
- 在HM3D和Gibson上实现强泛化,低质量轨迹也能学好探索。
- 三阶段训练提升空间理解,适合复杂环境导航任务。
本文提出MUVLA,一种面向物体导航的视觉-语言-动作地图理解模型。它通过语义地图抽象统一并结构化历史信息,以紧凑一致的形式编码空间上下文。MUVLA输入当前与历史观测及语义地图,根据目标物体描述预测动作序列。通过基于密集短时程进展信号的奖励引导回报建模,增强监督,使模型获得更细致的动作价值理解以最大化奖励。MUVLA采用三阶段训练:学习地图级空间理解、模仿混合质量示范行为、奖励放大。该策略将多样示范统一为鲁棒空间表示,生成更合理的探索策略。在HM3D和Gibson基准上的实验表明,MUVLA具备优异泛化能力,即使从低质量或部分成功轨迹中也能学习有效探索行为。
原文摘要 · Abstract (English)
In this paper, we present MUVLA, a Map Understanding Vision-Language-Action model tailored for object navigation. It leverages semantic map abstractions to unify and structure historical information, encoding spatial context in a compact and consistent form. MUVLA takes the current and history observations, as well as the semantic map, as inputs and predicts the action sequence based on the description of goal object. Furthermore, it amplifies supervision through reward-guided return modeling based on dense short-horizon progress signals, enabling the model to develop a detailed understanding of action value for reward maximization. MUVLA employs a three-stage training pipeline: learning map-level spatial understanding, imitating behaviors from mixed-quality demonstrations, and reward amplification. This strategy allows MUVLA to unify diverse demonstrations into a robust spatial representation and generate more rational exploration strategies. Experiments on HM3D and Gibson benchmarks demonstrate that MUVLA achieves great generalization and learns effective exploration behaviors even from low-quality or partially successful trajectories.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。