用视觉语言模型让机器人懂人情世故,实现安全社交导航
Vision-Language Models for Deployable Social Robot Navigation: Bridging Semantic Reasoning and Low-Level Control

- 构建三段式框架:高层语义推理、底层控制规划、中间桥梁机制
- 提出可部署的VLM导航路线图,融合语义理解与实际动作执行
- 适合研究社交机器人、具身智能与人机交互的开发者参考
社交机器人导航(SRN)不仅需要几何路径规划,还需理解人类意图、社会规范和情境线索,以生成符合社会规则的行为。传统导航方法虽能可靠规划路径并避障,但缺乏在复杂人机环境中的语义推理能力。近年来,视觉语言模型(VLMs)的发展为SRN带来了新机遇,支持高层语义理解、常识推理与自然语言交互。然而核心挑战在于如何将VLM的高层推理整合进实时、高安全性导航系统,并可靠地转化为具体的导航动作。本文提出一种统一视角,将现有方法归纳为三大组件:高层VLM推理、底层规划与控制、以及连接二者之间的中间机制。基于此,我们构建了结构化路线图,涵盖语义推理、评估器、空间对齐、中间表示与控制模块。该路线图既凸显了VLM的优势,也强调了混合架构对实际部署的必要性。此外,我们综述了代表性的数据集与评测平台,并讨论关键开放挑战。本综述旨在为构建可靠、社交合规且可部署的VLM增强型导航系统提供基础。
原文摘要 · Abstract (English)
Social robot navigation (SRN) requires more than geometric path planning; it demands understanding human intentions, social norms, and contextual cues to generate socially compliant behaviors. Although classical navigation methods provide reliable metric planning and collision avoidance, they often lack the semantic reasoning capabilities necessary for operation in complex human-centered environments. Recent advances in Vision-Language Models (VLMs) have opened new opportunities for SRN by enabling high-level VLM understanding, commonsense reasoning, and natural language interaction. However, a fundamental challenge remains: how to integrate VLMs into real-time, safety-critical navigation systems and reliably translate their high-level reasoning into grounded navigation actions. In this survey, we present a unified perspective of VLM-based SRN and organize existing approaches into three interconnected components: high-level VLM reasoning, low-level planning and control, and intermediate mechanisms that bridge reasoning and action. Based on this perspective, we propose a structured roadmap for coupling VLMs with navigation systems, covering semantic reasoning, evaluators, spatial grounding, intermediate representations, and control modules. The roadmap highlights both the strengths of VLMs and the necessity of hybrid architectures for practical deployment. We further review representative datasets and evaluation platforms developed for SRN. Finally, we discuss key open challenges. This survey aims to provide a foundation for building reliable, socially compliant, and deployable VLM-enabled navigation systems.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。