剖析强化学习导航中各模块贡献,提出性能提升新框架。
What Matters in RL-Based Methods for Object-Goal Navigation? An Empirical Study and A Unified Framework
- 拆解感知、策略、推理优化三模块,系统实验评估各自影响。
- 改进感知与推理策略,比单纯优化策略提升更显著,SPL增6.6%。
- 提供可复现框架与设计建议,适合机器人导航研究者参考。
物体目标导航(ObjectNav)是移动机器人在家庭、学校、办公等日常环境中部署的关键能力。该任务要求智能体仅依靠机载感知,在未见过的环境中定位指定类别物体实例,需融合语义理解、空间推理与长程规划。强化学习(RL)已成为主流方法,但现有系统在感知模块、策略架构和推理阶段策略上存在众多设计选择,其相对影响尚不明确。本文开展大规模实证研究,将导航流程分解为感知、策略与测试时增强三个核心模块,通过大量受控实验分析各模块独立贡献。结果表明,感知质量与测试时策略的改进常带来比单纯优化策略更大的性能提升,凸显模块间交互的重要性。基于此,我们提出统一框架以系统化研究模块化导航系统。据此构建的增强系统在Gibson基准上达到领先性能,成功率提升2.7%,SPL提高6.6%。此外,我们引入人类专家基线,成功率达98%,揭示当前RL智能体与人类水平之间的显著差距。最后,针对各模块提供实用洞见与设计建议,助力未来研究。
原文摘要 · Abstract (English)
Object-Goal Navigation (ObjectNav) is a key capability for deploying mobile robots in everyday environments such as homes, schools, and workplaces. In this task, an agent must locate an instance of a target object category in previously unseen environments using only onboard perception, requiring the integration of semantic understanding, spatial reasoning, and long-horizon planning. Reinforcement learning (RL) has become a dominant paradigm for ObjectNav, yet modern systems involve numerous design choices across perception modules, policy architectures, and inference-time strategies. The relative impact of these components, however, remains poorly understood. In this work, we present a large-scale empirical study of modular RL-based ObjectNav systems. We decompose the navigation pipeline into three key components: perception, policy, and test-time enhancement, and conduct extensive controlled experiments to analyze their individual contributions. Our results suggest that improvements in perception quality and test-time strategies often yield larger performance gains than policy improvements alone, highlighting the importance of understanding how different components interact within modular navigation systems. Motivated by these findings, we introduce a unified framework for systematically studying modular ObjectNav systems. Guided by our analysis, we build an enhanced system that achieves state-of-the-art performance on the Gibson benchmark, improving SPL by 6.6% and success rate by 2.7% over prior methods. We also introduce a human expert baseline, achieving 98% success, highlighting the significant gap between current RL agents and human-level navigation. Finally, we provide practical insights and design recommendations for each module to help guide future research. Project page: https://honwang0054.github.io/What-matters-in-RL-ObjNav-web/.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。