端到端生成情感丰富的对话虚拟人,实时响应且表情自然。
A$^2$-LLM: An End-to-end Conversational Audio Avatar Large Language Model
- 统一框架联合推理语言、语音韵律与3D面部动作
- 500毫秒延迟,0.7倍实时效率,情感表达更丰富
- 适合虚拟助手、游戏角色等需要真实互动的场景
构建富有表现力且响应迅速的对话式数字人是下一代人机交互的核心。尽管大语言模型显著提升了对话能力,当前多数系统仍依赖独立模块串联的级联架构,存在误差累积、高延迟和实时性差的问题。缺乏对对话上下文的理解,这类系统往往过度追求唇形同步,忽视情感深度。为此,我们提出A$^2$-LLM,一个端到端的对话音频虚拟人大语言模型,在统一框架内联合推理语言、音频韵律与3D面部运动。为支持训练,我们引入FLAME-QA,一个高质量多模态数据集,将语义意图与生动的面部动态在问答格式中对齐。通过深层语义理解,A$^2$-LLM生成的情感丰富面部动作超越简单唇同步。实验表明,本系统在保持实时效率(500毫秒延迟,0.7倍实时因子)的同时,实现了更优的情感表现力。
原文摘要 · Abstract (English)
Developing expressive and responsive conversational digital humans is a cornerstone of next-generation human-computer interaction. While large language models (LLMs) have significantly enhanced dialogue capabilities, most current systems still rely on cascaded architectures that connect independent modules. These pipelines are often plagued by accumulated errors, high latency, and poor real-time performance. Lacking access to the underlying conversational context, these pipelines inherently prioritize rigid lip-sync over emotional depth. To address these challenges, we propose A$^2$-LLM, an end-to-end conversational audio avatar large language model that jointly reasons about language, audio prosody, and 3D facial motion within a unified framework. To facilitate training, we introduce FLAME-QA, a high-quality multimodal dataset designed to align semantic intent with expressive facial dynamics within a QA format. By leveraging deep semantic understanding, A$^2$-LLM generates emotionally rich facial movements beyond simple lip-synchronization. Experimental results demonstrate that our system achieves superior emotional expressiveness while maintaining real-time efficiency (500 ms latency, 0.7 RTF).
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。