首个开源3D情感对话机器人,能实时读表情、说话带口型、渲染逼真形象。
EmpaAva: An Open-source Agentic 3D-Avatar Empathetic Live Chatbot

- 用三智能体架构实现语音/视觉情感感知与回应决策闭环
- 在情感理解与视听一致性上超越文本、2D头像等基线模型
- 适合研究人机交互、数字人开发或想快速搭建情感对话系统的开发者
本文提出EmpaAva,据我们所知首个开源的、具备代理能力的3D情感虚拟人实时对话系统。它将文本对话中的情感响应生成(ERG)拓展至面对面视频交互场景。通过类视频通话界面,用户与一个3D数字人对话,系统可从语音及可选视觉中识别情绪,并以带情感的语音、唇同步面部动作和照片级真实感3D高斯渲染进行回应。其核心为大语言模型驱动的三智能体架构:感知、共情响应规划与具身渲染形成闭环,配合响应规划层将每条回复转化为可执行的多模态指令,确保语音、表情与渲染统一于同一共情意图。基于强大的开源模块,EmpaAva提供可控制、可调试的整体智能体验。自动与人工评估均显示,EmpaAva在情感理解、回复质量与音视频一致性方面优于文本、2D说话头像及多模态虚拟人基线。我们已开源EmpaAva并提供在线实时演示。
原文摘要 · Abstract (English)
This paper presents EmpaAva, to our knowledge the first open-source, agentic 3D-avatar empathetic chatbot, which carries empathetic response generation (ERG) from text-only exchanges into live, face-to-face interaction. Through a video-call-like interface, a user speaks to a 3D digital human that reads their affect from speech and optional vision, and replies with emotional speech, lip-synced facial motion, and photorealistic 3D Gaussian rendering. At its core, an LLM coordinates a Tri-Agent Architecture, in which perception, empathetic response planning, and embodied rendering form a closed loop, paired with a Response Planning layer that compiles each reply into an executable multimodal plan, keeping voice, expression, and rendering on one empathetic intent. Building on strong open-source modules, EmpaAva supplies the intelligence that binds them into one controllable, inspectable experience. In automatic and human evaluations, EmpaAva surpasses text-only, 2D talking-face, and multimodal avatar baselines in emotion understanding, response quality, and audio-visual consistency. We open-source EmpaAva with an online live demo.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。