arXiv:2605.16816cs.RO2026-05

用视觉语言模型提升机器人对人情绪的理解,让协作更贴心。

"I'm Not Mad, Just Focused'': Understanding Human Emotions in Human-Robot Collaboration

论文配图:"I'm Not Mad, Just Focused'': Understanding Human Emotions in Human-Robot Collaboration
图 1 · 摘自论文原文
  • 用多模态视觉语言模型理解人类情绪,结合上下文提升判断力。
  • 相比传统模型,情绪识别与人工标注的语义和情感更一致。
  • 用户更喜欢能根据情绪调整行为的机器人,体验更自然。

人机协作(HRC)可受益于机器人对人类情绪状态的识别能力。然而,当前的HRC情绪识别(ER)模型常因依赖表演性数据集和单一模态输入(如面部表情)而表现不佳。本文提出一种基于视觉语言模型(VLM)的情绪识别系统,通过上下文理解提升情绪判读能力。首先,在现有HRC数据集上评估了该VLM-ER系统与人工标注在语义和情感相似性上的匹配度;随后,在一项服务机器人协同递送任务的用户研究中,测试了基于该系统推断用户情绪后调节机器人行为的效果。结果表明,所提VLM-ER系统在语义相似性和积极情感一致性上均优于基于卷积神经网络的基线模型。此外,用户更偏好由VLM-ER系统驱动的情绪自适应机器人行为。

原文摘要 · Abstract (English)

Human-robot collaboration (HRC) can benefit from robots' abilities to interpret human emotional states. However, current emotion recognition (ER) models in HRC often fall short, particularly due to their reliance on acted datasets and single-modality inputs like facial expressions. We propose a novel vision language model (VLM)-based ER system that leverages contextual understanding to improve emotion interpretation in HRC. We first evaluate the VLM-ER system by assessing its semantic and sentiment similarity with human annotations on an existing HRC dataset. Then, in a user study with a service robot in a collaborative delivery task, we evaluate the effects of modulating the robot's behaviour based on the user's emotional state inferred by the VLM-ER system. The results show that the proposed VLM-ER system achieves higher semantic similarity and positive sentiment alignment with human annotations compared to a baseline convolutional neural network-based system. Further, participants in the user study preferred emotion-adaptive robot behaviour facilitated by the VLM-ER system.

人机协作情绪识别视觉语言模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。