打造高保真实时对话数字人,支持自然互动与情感表达。
Hi-Reco: High-Fidelity Real-Time Conversational Digital Humans
- 异步流水线协调多模态组件,降低延迟实现实时响应。
- 引入历史增强与意图路由,提升对话连贯性与知识获取效率。
- 适合沉浸式通信、教育、娱乐场景,表现逼真且交互流畅。
高保真数字人日益应用于交互场景,但兼顾视觉真实感与实时响应仍是难题。本文提出一个高保真、实时对话数字人系统,无缝融合视觉逼真的3D形象、基于人格特征的语音合成以及基于知识的对话生成。为实现自然及时的交互,设计了异步执行流水线,最小化多模态组件间的延迟。系统支持唤醒词检测、情感化语调表达及高度准确、上下文感知的回应生成。通过新颖的检索增强方法,包括历史增强以维持对话连贯性,以及基于意图的路由实现高效知识访问。这些组件共同构成一个集成系统,使数字人具备强响应能力与可信表现,适用于通信、教育和娱乐等沉浸式应用。
原文摘要 · Abstract (English)
High-fidelity digital humans are increasingly used in interactive applications, yet achieving both visual realism and real-time responsiveness remains a major challenge. We present a high-fidelity, real-time conversational digital human system that seamlessly combines a visually realistic 3D avatar, persona-driven expressive speech synthesis, and knowledge-grounded dialogue generation. To support natural and timely interaction, we introduce an asynchronous execution pipeline that coordinates multi-modal components with minimal latency. The system supports advanced features such as wake word detection, emotionally expressive prosody, and highly accurate, context-aware response generation. It leverages novel retrieval-augmented methods, including history augmentation to maintain conversational flow and intent-based routing for efficient knowledge access. Together, these components form an integrated system that enables responsive and believable digital humans, suitable for immersive applications in communication, education, and entertainment.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。