arXiv:2409.20188cs.ROcs.SD2024-09被引 1

实时生成对话中听众的头部动作响应,提升人机互动自然度。

Active Listener: Continuous Generation of Listener's Head Motion Response in Dyadic Interactions

  • 基于图结构的端到端跨模态模型,直接从语音生成头部姿态。
  • 在IEMOCAP数据集上误差仅4.5度,帧率高适合实时应用。
  • 无需人工标注,能生成复杂连续动作,适用于人机交互场景。

对话中的非语言行为,如反映听者反应的头部动作,是双人言语互动的关键组成部分。尽管在生成伴随说话的动作方面已取得显著进展,但生成听者的响应动作仍是一大挑战。本文提出实时生成听者对说话者语音的连续头部动作响应的任务。为此,我们设计了一种基于图结构的端到端跨模态模型,以对话者的语音音频为输入,直接实时生成听者的头部姿态角(俯仰、翻滚、偏航)。与以往方法不同,本方法完全基于数据驱动,无需人工标注,也不将头部动作简化为简单的点头或摇头。在IEMOCAP数据集上的双人互动会话中进行的大量评估表明,该模型整体误差低至4.5度,且具有高帧率,具备在真实人机交互系统中部署的潜力。代码已公开于 https://github.com/bigzen/Active-Listener。

原文摘要 · Abstract (English)

A key component of dyadic spoken interactions is the contextually relevant non-verbal gestures, such as head movements that reflect a listener's response to the interlocutor's speech. Although significant progress has been made in the context of generating co-speech gestures, generating listener's response has remained a challenge. We introduce the task of generating continuous head motion response of a listener in response to the speaker's speech in real time. To this end, we propose a graph-based end-to-end crossmodal model that takes interlocutor's speech audio as input and directly generates head pose angles (roll, pitch, yaw) of the listener in real time. Different from previous work, our approach is completely data-driven, does not require manual annotations or oversimplify head motion to merely nods and shakes. Extensive evaluation on the dyadic interaction sessions on the IEMOCAP dataset shows that our model produces a low overall error (4.5 degrees) and a high frame rate, thereby indicating its deployability in real-world human-robot interaction systems. Our code is available at - https://github.com/bigzen/Active-Listener

动作生成人机交互语音理解实时系统

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。