arXiv:2605.31294cs.CV2026-05

用音频令牌直接生成实时生动人脸动画,告别机械感

TokTalk: Expressive Real-time Facial Animation from Audio-LLM Tokens

论文配图:TokTalk: Expressive Real-time Facial Animation from Audio-LLM Tokens
图 1 · 摘自论文原文
  • 直接从音频令牌生成面部动作,跳过中间文本环节
  • 实测延迟与现有方案相当,表现力和控制性显著更优
  • 适配任意语音大模型,支持聊天机器人等多场景应用

近期语音大模型(如 GPT-4o)的发展推动了语言模型的对话交互。然而,当前对话虚拟人仍存在表情僵硬、交流不自然的问题,主要源于语音识别、文本生成、分步响应、语音合成及音频驱动面部动画的多个串行步骤。我们发现,当前语音大模型输出的音频令牌已包含重建合理面部动作的充分信息,因此提出 TokTalk 系统,可直接从流式音频令牌实时生成富有表现力的 3D 面部动画。我们构建了一个新型的音频令牌到 3D 面部动作数据集,并基于分块条件流匹配模型进行训练。轻量级适配策略使模型能以极低计算开销无缝接入任意基于令牌的语音大模型。分块处理机制支持在延迟与面部质量间进行参数化权衡,消融实验验证了其有效性。感知评估显示,TokTalk 的实时性能在延迟上媲美现有方法,而在质量、表现力与控制性方面显著更优。我们通过聊天机器人虚拟人、语音驱动用户虚拟人及动画导演界面展示了系统的灵活性,适用于多样化的音视频人脸应用。

原文摘要 · Abstract (English)

Recent advances in Audio-LLMs like GPT-4o have ushered in an era of conversational interaction with language models. Conversational avatars however, still seem robotic in facial expression and conversational flow, in part due to sequential stages of speech recognition, text generation, turn-based text response, speech synthesis, and audio driven facial animation. Based on our insight that audio-tokens produced by current Audio-LLMs carry sufficient information to reconstruct a plausible facial performance, we present TokTalk, a system that directly outputs expressive facial animation in real-time from streaming audio-tokens. We construct a novel audio-token to 3D facial motion dataset, on which TokTalk is trained using a Chunk-based Conditional Flow Matching model. A lightweight adaptation strategy allows our trained model to seamlessly connect to any token-based Audio-LLM at minimal computational overhead. Our chunk-based processing further enables parametric trade-off between latency and facial quality, shown through ablation studies. We further show that the real-time performance of TokTalk is comparable in latency to prior art solutions, and significantly favorable (via a perceptual study) in terms of quality, expressivity and control of the 3D facial performance. We showcase TokTalk's flexibility using a chatbot Avatar, a voice-driven user Avatar, and an animation Director's interface, as diverse audio-visual face applications.

语音生成实时动画面部建模大模型应用

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。