用大模型提升机器人社交导航能力,减少对行人的干扰。
GSON: A Group-based Social Navigation Framework with Large Multimodal Model
- 基于视觉提示提取行人社交关系,实现零样本社交感知。
- 实测在排队、交谈等场景中显著降低对行人的干扰。
- 适合需要高社交智能的机器人导航应用。
随着服务机器人和自动驾驶车辆在人类环境中的普及,导航系统需从单纯抵达目的地演进为具备社会意识。本文提出GSON,一种基于群体的社交导航框架,利用大模态模型(LMMs)增强机器人的社会感知能力。该方法通过视觉提示实现行人社交关系的零样本提取,并与鲁棒的行人检测与跟踪流程结合,克服了LMM推理速度慢的瓶颈。规划系统采用介于全局路径规划与局部运动规划之间的中层规划器,有效兼顾全局上下文与实时响应性,同时避免对预测社交群体的破坏。我们在包含排队、对话和拍照等复杂社交场景的大量真实移动机器人导航实验中验证了GSON。对比结果表明,该系统在显著降低社交干扰的同时,保持了传统导航指标的相当性能。
原文摘要 · Abstract (English)
With the increasing presence of service robots and autonomous vehicles in human environments, navigation systems need to evolve beyond simple destination reach to incorporate social awareness. This paper introduces GSON, a novel group-based social navigation framework that leverages Large Multimodal Models (LMMs) to enhance robots' social perception capabilities. Our approach uses visual prompting to enable zero-shot extraction of social relationships among pedestrians and integrates these results with robust pedestrian detection and tracking pipelines to overcome the inherent inference speed limitations of LMMs. The planning system incorporates a mid-level planner that sits between global path planning and local motion planning, effectively preserving both global context and reactive responsiveness while avoiding disruption of the predicted social group. We validate GSON through extensive real-world mobile robot navigation experiments involving complex social scenarios such as queuing, conversations, and photo sessions. Comparative results show that our system significantly outperforms existing navigation approaches in minimizing social perturbations while maintaining comparable performance on traditional navigation metrics.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。