提出自适应双过程推理框架,提升视觉语言模型导航效率与成功率
Hydra-Nav: Object Navigation via Adaptive Dual-Process Reasoning
- 通过快慢系统切换实现高效决策与长期规划结合
- 在HM3D、MP3D、OVON上分别超越次优方法11.1%、17.4%、21.2%
- 引入新指标SOT,衡量不同推理强度下的搜索效率
尽管大型视觉语言模型(VLM)在物体目标导航中展现出潜力,但现有方法仍面临成功率低和未见物体定位效率差的问题,主要归因于薄弱的时空推理能力。近期尝试将推理注入VLM代理虽提升了成功率,但带来显著计算开销。为此,我们提出Hydra-Nav,一种统一的VLM架构,可自适应地在分析探索历史并制定高层计划的慢速思考系统与高效执行的快速反应系统之间切换。通过三阶段课程训练:(i) 空间-动作对齐以强化轨迹规划,(ii) 记忆-推理融合以增强长时程探索中的时空推理,(iii) 迭代拒绝微调以实现在关键决策点的选择性推理。大量实验表明,Hydra-Nav在HM3D、MP3D和OVON基准上达到当前最优性能,分别超越次优方法11.1%、17.4%和21.2%。此外,我们引入新指标SOT(Success weighted by Operation Time),用于衡量不同推理强度下VLM的搜索效率。结果表明,自适应推理显著优于固定频率基线。
原文摘要 · Abstract (English)
While large vision-language models (VLMs) show promise for object goal navigation, current methods still struggle with low success rates and inefficient localization of unseen objects--failures primarily attributed to weak temporal-spatial reasoning. Meanwhile, recent attempts to inject reasoning into VLM-based agents improve success rates but incur substantial computational overhead. To address both the ineffectiveness and inefficiency of existing approaches, we introduce Hydra-Nav, a unified VLM architecture that adaptively switches between a deliberative slow system for analyzing exploration history and formulating high-level plans, and a reactive fast system for efficient execution. We train Hydra-Nav through a three-stage curriculum: (i) spatial-action alignment to strengthen trajectory planning, (ii) memory-reasoning integration to enhance temporal-spatial reasoning over long-horizon exploration, and (iii) iterative rejection fine-tuning to enable selective reasoning at critical decision points. Extensive experiments demonstrate that Hydra-Nav achieves state-of-the-art performance on the HM3D, MP3D, and OVON benchmarks, outperforming the second-best methods by 11.1%, 17.4%, and 21.2%, respectively. Furthermore, we introduce SOT (Success weighted by Operation Time), a new metric to measure search efficiency across VLMs with varying reasoning intensity. Results show that adaptive reasoning significantly enhances search efficiency over fixed-frequency baselines.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。