arXiv:2506.02354cs.CV2025-06ACL被引 6

用视觉语言模型提升零样本物体导航的探索终止效率

RATE-Nav: Region-Aware Termination Enhancement for Zero-shot Object Navigation with Vision-Language Models

  • 基于区域感知设计探索终止机制,动态判断何时停止探索
  • 在HM3D上达67.8%成功率,MP3D上比前人方法提升约10%
  • 适合研究智能体探索策略与视觉语言模型应用的学者

物体导航(ObjectNav)是具身人工智能中的基础任务。尽管当前研究在语义地图构建和目标方向预测方面取得进展,冗余探索与探索失败仍不可避免。一个关键但未被充分探索的方向是及时终止探索以克服这些问题。我们观察到探索步数与探索率之间的边际效益递减,并分析了探索的成本收益关系。受此启发,提出RATE-Nav:一种区域感知的探索终止增强方法。包含几何预测区域分割算法与基于区域的探索估计算法,用于计算探索率。通过利用视觉语言模型(VLMs)的视觉问答能力,结合探索率实现高效终止。在HM3D数据集上取得67.8%的成功率和31.3%的SPL;在更具挑战性的MP3D数据集上,相比之前零样本方法提升约10%。

原文摘要 · Abstract (English)

Object Navigation (ObjectNav) is a fundamental task in embodied artificial intelligence. Although significant progress has been made in semantic map construction and target direction prediction in current research, redundant exploration and exploration failures remain inevitable. A critical but underexplored direction is the timely termination of exploration to overcome these challenges. We observe a diminishing marginal effect between exploration steps and exploration rates and analyze the cost-benefit relationship of exploration. Inspired by this, we propose RATE-Nav, a Region-Aware Termination-Enhanced method. It includes a geometric predictive region segmentation algorithm and region-Based exploration estimation algorithm for exploration rate calculation. By leveraging the visual question answering capabilities of visual language models (VLMs) and exploration rates enables efficient termination.RATE-Nav achieves a success rate of 67.8% and an SPL of 31.3% on the HM3D dataset. And on the more challenging MP3D dataset, RATE-Nav shows approximately 10% improvement over previous zero-shot methods.

物体导航视觉语言模型探索终止

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。