arXiv:2607.12329cs.HCcs.SD2026-07中稿 · 28th ACM Internati…

实时预测对话中听众点头的时机与运动参数,让虚拟人更自然地回应。

Real-time Generation of Listener Nodding via Prediction of Kinematic Parameters for Avatar Dialogue Systems

论文配图:Real-time Generation of Listener Nodding via Prediction of Kinematic Parameters for Avatar Dialogue Systems
图 1 · 摘自论文原文
  • 用双通道注意力网络基于语音活动投影,实时预测点头时机和运动特征。
  • 主观评测显示,相比随机或固定动作,新方法显著提升自然度。
  • 模型轻量高效,适合集成到实时对话虚拟人系统中。

在人类对话中,通过眼神交流、点头和面部表情等非语言线索并精确把握时机,实现流畅沟通。对话型虚拟人若能恰当表达这些线索,将更接近真实互动。本文聚焦于点头动作——其对展现主动倾听及鼓励用户继续发言至关重要。提出一种实时预测听众点头时机与运动学参数的模型,包含时间预测模块与运动参数预测模块,均基于语音活动投影(VAP)技术构建双通道注意力网络。与传统仅预测时间的方法不同,该模型可依据对话上下文动态生成具体运动特征。此外,验证了从已训练的时间预测模块初始化运动参数模块的微调有效性。所提模型轻量且支持实时运行,已集成至虚拟人对话系统。主观评价实验表明,本方法显著优于随机时间基线和固定动作基线。代码与训练模型已公开于 https://github.com/MaAI-Kyoto/MaAI。

原文摘要 · Abstract (English)

In human dialogue, we achieve smooth communication by expressing nonverbal cues such as eye contact, nodding, and facial expressions with precise timing. It is expected for conversational avatars to express these cues appropriately to realize natural and human-like interactions. This study focuses on nodding, which is crucial for demonstrating active listening and encouraging further user utterances. We propose a model that predicts both timing and kinematic parameters representing the motion features of listener nodding in real time. The proposed model consists of a timing prediction module and a kinematic parameter prediction module. Each implements a dyadic attention network over the speaker and listener channels based on the technique of Voice Activity Projection (VAP). Unlike conventional models, this approach enables real-time prediction of kinematic parameters based on the specific context of the dialogue rather than just predicting the timing. Furthermore, we demonstrate the effectiveness of fine-tuning the kinematic parameter prediction module initialized from the trained timing prediction module. The proposed model is lightweight and capable of real-time operation, and it has been integrated into an avatar dialogue system. Subjective evaluation experiments shows that our proposed method significantly outperforms both a baseline with stochastic timing and another with fixed-motion nodding. The code and trained models are available at https://github.com/MaAI-Kyoto/MaAI.

虚拟人对话系统实时生成点头模拟

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。