arXiv:2607.16956cs.RO2026-07

用视觉语言模型生成可解释的导航成本图,让机器人安全社交导航

G2-Nav: Grounded and Guarded Vision-Language Costmaps for Robot Social Navigation

论文配图:G2-Nav: Grounded and Guarded Vision-Language Costmaps for Robot Social Navigation
图 1 · 摘自论文原文
  • 将大模型语义推理转为可解释的视觉语言成本图
  • 实测在非结构化环境实现安全高效社交导航
  • 新增高频安全检测,应对系统延迟问题

社交导航需机器人在复杂现实环境中进行推理与响应。尽管近期工作尝试通过大型视觉-语言模型(VLM)融入人类级智能,但端到端框架常导致不可预测的黑箱行为,现有指令跟随方法亦未针对完全自主设计。为此,我们提出G2-Nav框架,实现抽象社交推理的具身化与安全部署保障。不直接让VLM做规划决策,而是将其语义推理转化为具备可靠性与可解释性的视觉-语言成本图。VLM评估开放集感知下的可通行区域与社会主体,将社交上下文映射至成本图。为提升真实场景鲁棒性,VLM对上游跟踪结果进行语义验证,并引入高频安全检查,在轨迹生成前防御系统延迟风险。通过真实实验表明,G2-Nav可在非结构化环境中实现安全、高效且符合社交规范的自主导航。代码将公开。

原文摘要 · Abstract (English)

Social navigation requires the robot to reason and respond in complex real-world environments. While recent works attempt to incorporate human-level intelligence into robot planning using large Vision-Language Models (VLMs), end-to-end frameworks often create an unpredictable black-box, and existing instruction-following methods are not designed for full autonomy. To bridge this gap, we present G2-Nav, a novel framework that grounds abstract social reasoning and guards safe real-world deployment. Instead of asking the VLM for direct planning decisions, G2-Nav translates its semantic reasoning into a vision-language costmap with reliability and interpretability. The VLM evaluates traversable regions and social agents from open-set perception, mapping social context into the costmap. To improve real-world robustness, the VLM performs semantic verification on upstream tracking, and we introduce a high-frequency safety check to guard against system latency prior to trajectory generation. We demonstrate through real-world experiments that G2-Nav delivers safe, efficient, and socially compliant autonomous navigation in unstructured environments. Code will be made publicly available.

机器人导航视觉语言模型成本图安全控制

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。