arXiv:2409.13675cs.RO2024-09ICRA被引 26

让机器人实时学习社交规则,边走边懂人情世故。

OLiVia-Nav: An Online Lifelong Vision Language Approach for Mobile Robot Social Navigation

  • 用轻量视觉语言模型实时理解社交与环境上下文
  • 在真实场景中降低路径误差和侵犯个人空间时长
  • 适合需长期适应新社交场景的服务机器人

服务机器人在医院、办公场所和养老院等以人为中心的环境中导航时,需遵守社交规范以保障人员安全与舒适。同时,还需适应导航过程中出现的新社交场景。本文提出一种新型在线终身视觉语言架构OLiVia-Nav,首次将视觉语言模型(VLMs)与在线终身学习框架结合,用于机器人社交导航。我们设计了一种独特的知识蒸馏方法——社交上下文对比语言图像预训练(SC-CLIP),将大型VLM的社会推理能力迁移到轻量级VLM中,使OLiVia-Nav能在导航时直接编码社交与环境上下文。这些嵌入用于生成并选择符合社交规范的机器人路径。SC-CLIP的终身学习能力使系统可随新社交场景的出现持续更新路径规划。我们在多种真实社交导航场景中进行了广泛实验。结果表明,OLiVia-Nav在均方误差、豪斯多夫损失和个人空间侵犯时长方面均优于现有最先进的深度强化学习与VLM方法。消融实验证实了其设计选择的有效性。

原文摘要 · Abstract (English)

Service robots in human-centered environments such as hospitals, office buildings, and long-term care homes need to navigate while adhering to social norms to ensure the safety and comfortability of the people they are sharing the space with. Furthermore, they need to adapt to new social scenarios that can arise during robot navigation. In this paper, we present a novel Online Lifelong Vision Language architecture, OLiVia- Nav, which uniquely integrates vision-language models (VLMs) with an online lifelong learning framework for robot social navigation. We introduce a unique distillation approach, Social Context Contrastive Language Image Pre-training (SC-CLIP), to transfer the social reasoning capabilities of large VLMs to a lightweight VLM, in order for OLiVia-Nav to directly encode social and environment context during robot navigation. These encoded embeddings are used to generate and select robot social compliant trajectories. The lifelong learning capabilities of SC-CLIP enable OLiVia-Nav to update the robot trajectory planning overtime as new social scenarios are encountered. We conducted extensive real-world experiments in diverse social navigation scenarios. The results showed that OLiVia-Nav outperformed existing state-of-the-art DRL and VLM methods in terms of mean squared error, Hausdorff loss, and personal space violation duration. Ablation studies also verified the design choices for OLiVia-Nav.

机器人导航视觉语言模型终身学习社会交互

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。