arXiv:2501.09024cs.CVcs.HC2025-01被引 26

用语言模型让机器人像人一样在人群中自然导航。

Social-LLaVA: Enhancing Robot Navigation through Human-Language Reasoning in Social Spaces

  • 用视觉语言模型连接感知与社交动作,实现类人推理
  • 在50个问答上平均得分超越GPT-4V和Gemini
  • 适合研究社交机器人、人机交互与具身智能的学者

现有社交机器人导航方法多依赖手工规则或人类示范,难以有效将感知转化为符合社交规范的动作。为弥合这一差距,我们借鉴视觉语言模型(VLM)的成功,提出通过语言实现感知与社交行为间的类人推理。为此构建了名为SNEI的视觉语言数据集,包含40,000条基于2,000次真实人-机器人社交互动的视觉问答(VQA),覆盖感知、预测、思维链推理、行动与解释等环节。我们使用SNEI微调出Social-LLaVA模型,在50个视觉问答任务中,经15项人工评分平均表现优于GPT-4V和Gemini。将其部署于移动机器人上,实现了接近人类的动态公共空间导航能力,标志着语言驱动社交导航的重要进展。

原文摘要 · Abstract (English)

Most existing social robot navigation techniques either leverage hand-crafted rules or human demonstrations to connect robot perception to socially compliant actions. However, there remains a significant gap in effectively translating perception into socially compliant actions, much like how human reasoning naturally occurs in dynamic environments. Considering the recent success of Vision-Language Models (VLMs), we propose using language to bridge the gap in human-like reasoning between perception and socially aware robot actions. We create a vision-language dataset, Social robot Navigation via Explainable Interactions (SNEI), featuring 40K human-annotated Visual Question Answers (VQAs) based on 2K human-robot social interactions in unstructured, crowded public spaces, spanning perception, prediction, chain-of-thought reasoning, action, and explanation. We fine-tune a VLM, Social-LLaVA, using SNEI to demonstrate the practical application of our dataset. Social-LLaVA outperforms state-of-the-art models like GPT-4V and Gemini, based on the average of fifteen different human-judge scores across 50 VQA. Deployed onboard a mobile robot, Social-LLaVA enables human-like reasoning, marking a promising step toward socially compliant robot navigation in dynamic public spaces through language reasoning.

社交机器人视觉语言模型人机交互导航

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。