评测视觉语言模型在社交机器人导航中的场景理解能力。
SocialNav-SUB: Benchmarking VLMs for Scene Understanding in Social Robot Navigation
- 构建VQA基准数据集SocialNav-SUB,评估模型对社交场景的多维理解。
- 顶尖VLM仍低于规则基线和人类共识,暴露社会场景理解短板。
- 适合研究社交机器人、视觉语言模型与人机交互的学者参考。
动态、以人类为中心的环境中机器人导航需要基于可靠场景理解的社会合规决策。近年来的视觉语言模型(VLMs)展现出物体识别、常识推理和上下文理解等潜力,契合社交机器人导航的复杂需求。然而,现有研究尚未系统评估其在复杂社交导航场景中(如推断代理间时空关系及人类意图)的理解能力,而这对于安全、合乎社会规范的导航至关重要。本文提出社会导航场景理解基准(SocialNav-SUB),一个面向真实社交机器人导航场景的视觉问答(VQA)数据集与评测框架,用于统一评估VLMs在空间、时空与社会推理任务中的表现,并对比人类与规则基线。实验表明,尽管最优VLM与人类答案有较高一致性,但仍显著落后于简单规则方法与人类共识,揭示当前VLM在社会场景理解方面存在关键差距。该基准为面向社交机器人导航的基础模型研究提供新方向,助力探索如何适配VLM以满足实际需求。代码与数据见:https://larg.github.io/socialnav-sub。
原文摘要 · Abstract (English)
Robot navigation in dynamic, human-centered environments requires socially-compliant decisions grounded in robust scene understanding. Recent Vision-Language Models (VLMs) exhibit promising capabilities such as object recognition, common-sense reasoning, and contextual understanding-capabilities that align with the nuanced requirements of social robot navigation. However, it remains unclear whether VLMs can accurately understand complex social navigation scenes (e.g., inferring the spatial-temporal relations among agents and human intentions), which is essential for safe and socially compliant robot navigation. While some recent works have explored the use of VLMs in social robot navigation, no existing work systematically evaluates their ability to meet these necessary conditions. In this paper, we introduce the Social Navigation Scene Understanding Benchmark (SocialNav-SUB), a Visual Question Answering (VQA) dataset and benchmark designed to evaluate VLMs for scene understanding in real-world social robot navigation scenarios. SocialNav-SUB provides a unified framework for evaluating VLMs against human and rule-based baselines across VQA tasks requiring spatial, spatiotemporal, and social reasoning in social robot navigation. Through experiments with state-of-the-art VLMs, we find that while the best-performing VLM achieves an encouraging probability of agreeing with human answers, it still underperforms simpler rule-based approach and human consensus baselines, indicating critical gaps in social scene understanding of current VLMs. Our benchmark sets the stage for further research on foundation models for social robot navigation, offering a framework to explore how VLMs can be tailored to meet real-world social robot navigation needs. An overview of this paper along with the code and data can be found at https://larg.github.io/socialnav-sub .
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。