让虚拟听众头动更生动可控,支持长序列情感表达与多模态交互。
VividListener: Expressive and Controllable Listener Dynamics Modeling for Multi-Modal Responsive Interaction
- 通过多模态条件引导,实现听者动作与语义、情绪的精细对齐。
- 在140万帧的ListenerX数据集上达到当前最佳表现,支持情绪强度可调。
- 适合虚拟人对话系统、动画生成等需要自然反应的应用场景。
生成具有细腻情感和表现力的听者头部动态对于虚拟角色对话建模至关重要。以往研究多聚焦短期行为生成,忽视了长期序列中运动变化与情绪强度的精细控制。同时,缺乏包含头部动作与细粒度多模态标注(如文本描述、情绪强度)的大型配对语料库也限制了发展。为此,我们首次构建了一个大规模多轮3D双人对话数据集ListenerX,包含超过140万有效帧,用于多模态响应交互建模。并提出VividListener框架,实现可精细控制、富有表现力的听者动态建模。该框架利用多模态条件作为引导,设计了响应交互模块(RIM)以自适应表示多模态交互嵌入,确保听者动作与文本描述及说话人行为在语义和表现上保持一致。同时引入情绪强度标签(EIT),融合多模态信息实现情绪强度编辑,适用于文本描述与动作幅度调节。在新构建的ListenerX数据集上的大量实验表明,VividListener在表达性与可控性上均达到先进水平。
原文摘要 · Abstract (English)
Generating responsive listener head dynamics with nuanced emotions and expressive reactions is crucial for practical dialogue modeling in various virtual avatar animations. Previous studies mainly focus on the direct short-term production of listener behavior. They overlook the fine-grained control over motion variations and emotional intensity, especially in long-sequence modeling. Moreover, the lack of long-term and large-scale paired speaker-listener corpora including head dynamics and fine-grained multi-modality annotations (e.g., text-based expression descriptions, emotional intensity) also limits the application of dialogue modeling.Therefore, we first newly collect a large-scale multi-turn dataset of 3D dyadic conversation containing more than 1.4M valid frames for multi-modal responsive interaction, dubbed ListenerX. Additionally, we propose VividListener, a novel framework enabling fine-grained, expressive and controllable listener dynamics modeling. This framework leverages multi-modal conditions as guiding principles for fostering coherent interactions between speakers and listeners.Specifically, we design the Responsive Interaction Module (RIM) to adaptively represent the multi-modal interactive embeddings. RIM ensures the listener dynamics achieve fine-grained semantic coordination with textual descriptions and adjustments, while preserving expressive reaction with speaker behavior. Meanwhile, we design the Emotional Intensity Tags (EIT) for emotion intensity editing with multi-modal information integration, applying to both text descriptions and listener motion amplitude.Extensive experiments conducted on our newly collected ListenerX dataset demonstrate that VividListener achieves state-of-the-art performance, realizing expressive and controllable listener dynamics.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。