评测视觉语言模型在人机交互中的实时感知能力,发现现有模型性能与延迟难兼顾。
HRIBench: Benchmarking Vision-Language Models for Real-Time Human Perception in Human-Robot Interaction
- 构建包含1000个问题的多领域测评基准,覆盖非语言信号理解等5类核心人机交互任务。
- 11个主流模型在真实场景下表现不佳,平均准确率不足60%,且响应延迟过高。
- 适合关注机器人实时感知、模型效率优化的研究者和开发者参考。
实时人类感知对高效人机交互至关重要。大型视觉语言模型(VLM)虽具备良好泛化能力,但常因高延迟影响用户体验,限制其在真实场景的应用。为系统评估VLM在人机交互中的人类感知能力及性能-延迟权衡,我们提出了HRIBench,一个面向人机交互关键感知任务的视觉问答(VQA)基准。该基准涵盖五个核心领域:(1) 非语言线索理解,(2) 口语指令理解,(3) 人-机物体关系理解,(4) 社交导航,(5) 人物识别。数据来自真实人机交互环境,非语言线索部分为自建,其余四领域使用公开数据集。每个领域精心设计200个问题,共1000个问题。我们对11个顶尖闭源与开源VLM进行了全面评估。结果显示,尽管具备泛化能力,当前模型在核心感知任务上仍表现欠佳;所有模型均未达到适用于实时部署的性能-延迟平衡,凸显未来需研发更小、低延迟、强感知能力的VLM。HRIBench与实验结果可在GitHub获取:https://github.com/interaction-lab/HRIBench。
原文摘要 · Abstract (English)
Real-time human perception is crucial for effective human-robot interaction (HRI). Large vision-language models (VLMs) offer promising generalizable perceptual capabilities but often suffer from high latency, which negatively impacts user experience and limits VLM applicability in real-world scenarios. To systematically study VLM capabilities in human perception for HRI and performance-latency trade-offs, we introduce HRIBench, a visual question-answering (VQA) benchmark designed to evaluate VLMs across a diverse set of human perceptual tasks critical for HRI. HRIBench covers five key domains: (1) non-verbal cue understanding, (2) verbal instruction understanding, (3) human-robot object relationship understanding, (4) social navigation, and (5) person identification. To construct HRIBench, we collected data from real-world HRI environments to curate questions for non-verbal cue understanding, and leveraged publicly available datasets for the remaining four domains. We curated 200 VQA questions for each domain, resulting in a total of 1000 questions for HRIBench. We then conducted a comprehensive evaluation of both state-of-the-art closed-source and open-source VLMs (N=11) on HRIBench. Our results show that, despite their generalizability, current VLMs still struggle with core perceptual capabilities essential for HRI. Moreover, none of the models within our experiments demonstrated a satisfactory performance-latency trade-off suitable for real-time deployment, underscoring the need for future research on developing smaller, low-latency VLMs with improved human perception capabilities. HRIBench and our results can be found in this Github repository: https://github.com/interaction-lab/HRIBench.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。