arXiv:2410.18373cs.ROcs.HC2024-10ICRA被引 8

让机器人在多人对话中精准识情表意,实时响应且抗干扰。

UGotMe: An Embodied System for Affective Human-Robot Interaction

  • 用主动人脸提取过滤无关物体和静止说话者,减少视觉噪声。
  • 通过高效数据传输实现机器人端实时推理,延迟可控。
  • 专为真实多人互动场景设计,适合需要情感交互的机器人应用。

赋予人形机器人理解人类情绪并根据情境恰当表达情绪的能力,是实现情感化人机交互的关键。然而,将现有视觉感知的多模态情绪识别模型应用于真实世界的人机交互,面临两大挑战:环境噪声干扰与实时性要求。首先,在多人对话场景中,机器人视觉输入受干扰因素影响,包括场景中的分散物体或视野内的非活跃说话者,阻碍情绪线索提取。其次,实时响应难以实现。为此,我们提出名为UGotMe的情感化人机交互系统,专为多人对话设计。系统引入两种去噪策略:一是从原始图像中提取说话者人脸,二是采用定制化主动人脸提取机制,排除非活跃说话者。针对实时性问题,通过优化机器人与本地服务器间的数据传输流程提升响应效率。我们在人形机器人Ameca上部署UGotMe,验证其在实际场景中的实时推理能力。演示视频详见 https://lipzh5.github.io/HumanoidVLE/。

原文摘要 · Abstract (English)

Equipping humanoid robots with the capability to understand emotional states of human interactants and express emotions appropriately according to situations is essential for affective human-robot interaction. However, enabling current vision-aware multimodal emotion recognition models for affective human-robot interaction in the real-world raises embodiment challenges: addressing the environmental noise issue and meeting real-time requirements. First, in multiparty conversation scenarios, the noises inherited in the visual observation of the robot, which may come from either 1) distracting objects in the scene or 2) inactive speakers appearing in the field of view of the robot, hinder the models from extracting emotional cues from vision inputs. Secondly, realtime response, a desired feature for an interactive system, is also challenging to achieve. To tackle both challenges, we introduce an affective human-robot interaction system called UGotMe designed specifically for multiparty conversations. Two denoising strategies are proposed and incorporated into the system to solve the first issue. Specifically, to filter out distracting objects in the scene, we propose extracting face images of the speakers from the raw images and introduce a customized active face extraction strategy to rule out inactive speakers. As for the second issue, we employ efficient data transmission from the robot to the local server to improve realtime response capability. We deploy UGotMe on a human robot named Ameca to validate its real-time inference capabilities in practical scenarios. Videos demonstrating real-world deployment are available at https://lipzh5.github.io/HumanoidVLE/.

情感交互人形机器人实时系统多模态

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。