让机器人听懂指令,在复杂社交场景中安全导航
LISN: Language-Instructed Social Navigation with VLM-based Controller Modulating
- 用视觉语言模型动态调节导航参数,实现快速响应与精准避障
- 91.3%成功率,显著优于基线(高出63%),尤其在人群跟随和禁入区规避中表现突出
- 首个支持指令理解的社交导航仿真基准,适合人机共融研究者
为实现人机共存,社交感知导航对移动机器人至关重要。现有研究多聚焦路径效率与行人碰撞规避,但仅涵盖社交导航的一部分。机器人还需遵从用户指令,使行为符合任务目标与人类社会规范。本文提出LISN-Bench,首个基于仿真的语言指令社交导航基准,基于Rosnav-Arena 3.0构建,首次在多样化场景中整合指令遵循与场景理解。为此,我们提出Social-Nav-Modulator,一种快慢分层系统:视觉语言模型(VLM)代理动态调节代价地图与控制器参数。将底层动作生成与较慢的VLM推理解耦,降低高频VLM调用依赖,同时提升动态避障与感知适应能力。方法在挑战性任务中表现优异,平均成功率达91.3%,相比最优基线提升超过63%。
原文摘要 · Abstract (English)
Towards human-robot coexistence, socially aware navigation is significant for mobile robots. Yet existing studies on this area focus mainly on path efficiency and pedestrian collision avoidance, which are essential but represent only a fraction of social navigation. Beyond these basics, robots must also comply with user instructions, aligning their actions to task goals and social norms expressed by humans. In this work, we present LISN-Bench, the first simulation-based benchmark for language-instructed social navigation. Built on Rosnav-Arena 3.0, it is the first standardized social navigation benchmark to incorporate instruction following and scene understanding across diverse contexts. To address this task, we further propose Social-Nav-Modulator, a fast-slow hierarchical system where a VLM agent modulates costmaps and controller parameters. Decoupling low-level action generation from the slower VLM loop reduces reliance on high-frequency VLM inference while improving dynamic avoidance and perception adaptability. Our method achieves an average success rate of 91.3%, which is greater than 63% than the most competitive baseline, with most of the improvements observed in challenging tasks such as following a person in a crowd and navigating while strictly avoiding instruction-forbidden regions. The project website is at: https://social-nav.github.io/LISN-project/
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。