arXiv:2507.23298cs.HCcs.SD2025-07中稿 · 27th ACM Internati…被引 4

实时预测不同类型的点头动作,让虚拟角色更自然地回应对话

Real-time Generation of Various Types of Nodding for Avatar Attentive Listening System

  • 基于语音活动投影模型,实时预测点头时机与类型
  • 多任务学习提升准确率,低处理速率仍保持高精度
  • 适合虚拟助手、人机交互系统等需要自然非语言反馈的场景

在人类对话中,点头和面部表情等非语言信息与语言信息同样重要,对话系统也应具备此类表现能力。本文聚焦于关注性倾听系统中的点头行为,提出一种可实时预测其时机与类型的模型。该模型基于语音活动投影(VAP)模型,从说话人和听者音频中预测语音活动,并扩展为连续、实时地生成多种类型点头。同时引入多任务学习,联合预测口语回应信号,并在通用对话数据上进行预训练。实验表明,多任务学习显著提升了定时与类型预测效果。通过降低处理频率,模型实现低延迟实时运行,且准确率下降有限,已集成至虚拟角色倾听系统中。主观评估显示,该方法优于始终与口头回应同步的常规方法。代码与训练模型已在 https://github.com/MaAI-Kyoto/MaAI 开源。

原文摘要 · Abstract (English)

In human dialogue, nonverbal information such as nodding and facial expressions is as crucial as verbal information, and spoken dialogue systems are also expected to express such nonverbal behaviors. We focus on nodding, which is critical in an attentive listening system, and propose a model that predicts both its timing and type in real time. The proposed model builds on the voice activity projection (VAP) model, which predicts voice activity from both listener and speaker audio. We extend it to prediction of various types of nodding in a continuous and real-time manner unlike conventional models. In addition, the proposed model incorporates multi-task learning with verbal backchannel prediction and pretraining on general dialogue data. In the timing and type prediction task, the effectiveness of multi-task learning was significantly demonstrated. We confirmed that reducing the processing rate enables real-time operation without a substantial drop in accuracy, and integrated the model into an avatar attentive listening system. Subjective evaluations showed that it outperformed the conventional method, which always does nodding in sync with verbal backchannel. The code and trained models are available at https://github.com/MaAI-Kyoto/MaAI.

虚拟角色非语言交互实时生成

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。