让机器人通过视觉语言模型实现多模态社交对话
Towards Multimodal Social Conversations with Robots: Using Vision-Language Models
- 用视觉语言模型融合图像与文本信息理解社交场景
- 突破传统仅靠语音对话的局限,支持环境感知互动
- 适合研究人机交互、具身智能的学者和开发者
大语言模型使社交机器人具备了自主开展开放域对话的能力。然而,它们仍缺乏关键的社会技能:利用多种模态传递社交信息。以往工作集中于需要参考环境或特定现象(如对话中断)的任务导向交互,而本文则概述了机器人进行多模态社交对话的整体需求。我们提出,视觉语言模型能够以足够通用的方式处理广泛的视觉信息,适用于自主社交机器人。文中描述了如何适配此类模型,并讨论了现存的技术挑战及评估方法。
原文摘要 · Abstract (English)
Large language models have given social robots the ability to autonomously engage in open-domain conversations. However, they are still missing a fundamental social skill: making use of the multiple modalities that carry social interactions. While previous work has focused on task-oriented interactions that require referencing the environment or specific phenomena in social interactions such as dialogue breakdowns, we outline the overall needs of a multimodal system for social conversations with robots. We then argue that vision-language models are able to process this wide range of visual information in a sufficiently general manner for autonomous social robots. We describe how to adapt them to this setting, which technical challenges remain, and briefly discuss evaluation practices.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。