实时生成自然多样肢体表情,让对话机器人更像真人。
MIBURI: Towards Expressive Interactive Gesture Synthesis
- 分部位编码动作,用二维因果模型实时生成
- 生成动作自然且贴合语境,避免僵硬重复
- 适合需要实时交互的虚拟人、聊天机器人场景
具身对话代理(ECAs)旨在通过语音、手势和面部表情模拟面对面交流。现有基于大语言模型(LLM)的对话代理缺乏身体表现力,而传统动作生成方法或动作呆板、多样性低,或依赖未来语音信息且运行时间长。为此,我们提出MIBURI,首个在线、因果的全身体态与面部表情生成框架,可实时同步语音对话生成富有表现力的动作。采用分部位感知的动作编码器,将层级化运动细节编码为多层离散标记,并通过二维因果模型,基于LLM生成的语音文本嵌入,实时自回归生成动作,同时建模时间动态与部件级运动层次。引入辅助目标以增强表现力与多样性,防止陷入静态姿态。对比实验表明,该方法在自然性与语境对齐性上优于近期基线。更多效果请参见演示视频:https://vcai.mpi-inf.mpg.de/projects/MIBURI/
原文摘要 · Abstract (English)
Embodied Conversational Agents (ECAs) aim to emulate human face-to-face interaction through speech, gestures, and facial expressions. Current large language model (LLM)-based conversational agents lack embodiment and the expressive gestures essential for natural interaction. Existing solutions for ECAs often produce rigid, low-diversity motions, that are unsuitable for human-like interaction. Alternatively, generative methods for co-speech gesture synthesis yield natural body gestures but depend on future speech context and require long run-times. To bridge this gap, we present MIBURI, the first online, causal framework for generating expressive full-body gestures and facial expressions synchronized with real-time spoken dialogue. We employ body-part aware gesture codecs that encode hierarchical motion details into multi-level discrete tokens. These tokens are then autoregressively generated by a two-dimensional causal framework conditioned on LLM-based speech-text embeddings, modeling both temporal dynamics and part-level motion hierarchy in real time. Further, we introduce auxiliary objectives to encourage expressive and diverse gestures while preventing convergence to static poses. Comparative evaluations demonstrate that our causal and real-time approach produces natural and contextually aligned gestures against recent baselines. We urge the reader to explore demo videos on https://vcai.mpi-inf.mpg.de/projects/MIBURI/.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。