让虚拟听众根据讲话内容生成自然反应动作
ReactMotion: Generating Reactive Listener Motions from Speaker Utterance
- 联合建模文本、音频、情绪与动作,生成多样化反应
- 在多候选动作数据集上训练,提升反应恰当性
- 适合虚拟人交互、动画生成等需要自然非语言行为的场景
本文提出一项新任务:从讲话内容生成自然的听众身体反应动作。由于人类反应具有内在不确定性,该任务长期缺乏研究。为此,我们构建了ReactMotionNet——一个大规模数据集,将讲话内容与多个不同恰当度的听众动作配对,明确捕捉动作的“一因多果”特性,并提供超越单一真值的动作监督。基于此,我们设计了面向偏好的评估协议,专门衡量反应的恰当性。进一步提出ReactMotion框架,统一建模文本、音频、情绪与动作,采用偏好学习目标训练,鼓励生成既恰当又多样的反应动作。大量实验表明,该方法优于检索基线和级联式大模型流程,生成的动作更自然、多样且贴切。
原文摘要 · Abstract (English)
In this paper, we introduce a new task, Reactive Listener Motion Generation from Speaker Utterance, which aims to generate naturalistic listener body motions that appropriately respond to a speaker's utterance. However, modeling such nonverbal listener behaviors remains underexplored and challenging due to the inherently non-deterministic nature of human reactions. To facilitate this task, we present ReactMotionNet, a large-scale dataset that pairs speaker utterances with multiple candidate listener motions annotated with varying degrees of appropriateness. This dataset design explicitly captures the one-to-many nature of listener behavior and provides supervision beyond a single ground-truth motion. Building on this dataset design, we develop preference-oriented evaluation protocols tailored to evaluate reactive appropriateness, where conventional motion metrics focusing on input-motion alignment ignore. We further propose ReactMotion, a unified generative framework that jointly models text, audio, emotion, and motion, and is trained with preference-based objectives to encourage both appropriate and diverse listener responses. Extensive experiments show that ReactMotion outperforms retrieval baselines and cascaded LLM-based pipelines, generating more natural, diverse, and appropriate listener motions.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。