让机器人导航时懂礼节,不只避障还顾及社交规范。
From Obstacles to Etiquette: Robot Social Navigation with VLM-Informed Path Selection
- 用视觉语言模型评估路径的社交合理性
- 实测显示违规时间最短且无社交区入侵
- 适合需与人近距离互动的智能机器人
在人类环境中进行社会性导航不仅需满足几何约束,因即使无碰撞的路径仍可能干扰正在进行的活动或违背社交规范。为此,需分析代理间交互并融入常识推理。本文提出一种融合几何规划与上下文社交推理的机器人导航框架:先提取障碍物与人体动态生成几何可行路径候选,再通过微调的视觉语言模型(VLM)基于情境化社会预期评估路径,选择最优社交路径供控制器执行。该任务特定的VLM将大模型中的社会推理提炼为小型高效模型,支持在多样化人机交互场景中实时适应。四个社交导航场景的实验表明,本方法整体表现最佳,个人空间侵犯时长最低,行人面对时间最少,且无社交区域入侵。
原文摘要 · Abstract (English)
Navigating socially in human environments requires more than satisfying geometric constraints, as collision-free paths may still interfere with ongoing activities or conflict with social norms. Addressing this challenge calls for analyzing interactions between agents and incorporating common-sense reasoning into planning. This paper presents a social robot navigation framework that integrates geometric planning with contextual social reasoning. The system first extracts obstacles and human dynamics to generate geometrically feasible candidate paths, then leverages a fine-tuned vision-language model (VLM) to evaluate these paths, informed by contextually grounded social expectations, selecting a socially optimized path for the controller. This task-specific VLM distills social reasoning from large foundation models into a smaller and efficient model, allowing the framework to perform real-time adaptation in diverse human-robot interaction contexts. Experiments in four social navigation contexts demonstrate that our method achieves the best overall performance with the lowest personal space violation duration, the minimal pedestrian-facing time, and no social zone intrusions. Project page: https://path-etiquette.github.io
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。