arXiv:2503.15491cs.HCcs.CL2025-03被引 3

用大模型判断机器人何时该主动互动,提升人机交互启动效率。

Agreeing to Interact in Human-Robot Interaction using Large Language Models and Vision Language Models

  • 结合大语言模型与视觉语言模型判断互动时机
  • 在明确动作场景下准确率达GPT-4o表现优异
  • 开放场景中仍需平衡人机状态,挑战较大

在人机交互(HRI)中,互动开始阶段往往复杂。机器人是否应与人类交流,取决于多种情境因素(如当前人类活动、交互紧迫性等)。本文测试大语言模型(LLM)和视觉语言模型(VLM)能否解决此问题。对比四种基于LLM和VLM的系统设计模式,在包含84个人机交互场景的测试集上进行评估。该测试集融合多个公开数据集,并包含若干开放性决策场景。结果表明,使用GPT-4o和Phi-3 Vision模型时,LLM和VLM在目标动作明确的情况下能有效处理互动起始;但在开放性场景中,模型需在人类与机器人状态间权衡,仍面临挑战。

原文摘要 · Abstract (English)

In human-robot interaction (HRI), the beginning of an interaction is often complex. Whether the robot should communicate with the human is dependent on several situational factors (e.g., the current human's activity, urgency of the interaction, etc.). We test whether large language models (LLM) and vision language models (VLM) can provide solutions to this problem. We compare four different system-design patterns using LLMs and VLMs, and test on a test set containing 84 human-robot situations. The test set mixes several publicly available datasets and also includes situations where the appropriate action to take is open-ended. Our results using the GPT-4o and Phi-3 Vision model indicate that LLMs and VLMs are capable of handling interaction beginnings when the desired actions are clear, however, challenge remains in the open-ended situations where the model must balance between the human and robot situation.

人机交互大模型视觉语言模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。