arXiv:2607.16318cs.HCcs.RO2026-07

用视觉语言模型让机器人理解人类互动情境,提升对话自然度。

The World According to a Social Robot -- Augmenting Human-Robot Dialogue With Vision Language Models

论文配图:The World According to a Social Robot -- Augmenting Human-Robot Dialogue With Vision Language Models
图 1 · 摘自论文原文
  • 将视觉语言模型接入Pepper机器人,实现跨模态对话理解。
  • 视觉信息使对话更贴合场景,响应时间仅小幅增加。
  • 欧洲本地部署大模型,符合数据合规要求,适合真实应用。

视觉语言模型(VLMs)使机器人能够感知环境及对话者的行为与特征,尤其在日常场景中部署社交机器人时,理解符合人类习惯的情境至关重要。本文探讨了将Mistral AI语言模型与Pepper机器人结合用于人机对话的初步经验,并研究了额外视觉信息对不同模型响应时间的影响。结果显示,引入视觉信息可显著增强对话上下文,使机器人与人类能共同关注未言明的情境要素,同时响应时间仅略有增加。此外,使用欧洲本地托管的大语言模型(LLM)满足欧洲数据保护法规,有助于推动实际应用场景落地。

原文摘要 · Abstract (English)

Vision Language Models (VLMs) enable robots to visually perceive their environment as well as the actions and characteristics of their conversation partner or humans in collaboration. Especially for social robots deployed in everyday settings and for uncomplicated, natural use, it is essential that the robot has an understanding of situations that is appropriate to human customs. This paper presents initial experiences with the application of a Mistral AI language model with a Pepper robot for Human-Robot Interaction (HRI) in dialogue, as well as an investigation of the effects of additional visual information on response time in different models. The results show that incorporating visual information adds context to the dialogue with only a moderate increase in response time, enabling both the robot and the human to take into account unspoken elements of the situation. Furthermore, using an LLM hosted in Europe offers a solution that complies with European data protection regulations and can therefore facilitate real-life applications more easily.

人机交互视觉语言模型社交机器人多模态

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。