arXiv:2605.28272cs.CV2026-05International Conf…被引 1

实时驱动3D虚拟人动作,支持语音与音乐同步生成

EchoAvatar: Real-time Generative Avatar Animation from Audio Streams

论文配图:EchoAvatar: Real-time Generative Avatar Animation from Audio Streams
图 1 · 摘自论文原文
  • 统一流式架构,逐段输入音频即时生成连贯动作
  • 在真实对话和音乐场景下均实现高保真动作同步
  • 支持大模型指令控制,适合交互式虚拟助手应用

实时从音频流生成高保真3D角色动作是下一代互动虚拟人和虚拟助手的关键。然而,现有方法多限于离线处理完整音频序列或仅适用于特定领域,难以同时有效处理语音与音乐。本文提出一种新框架,可从持续输入的语音和音乐中生成连续、连贯的全身动作,延迟极低。核心是统一的流式架构,能基于增量音频输入合成连续动作。采用强音频依赖训练策略,使模型无需显式域标签即可无缝泛化至对话语音与节奏性音乐。此外,探索使用强化学习优化在线生成质量。通过工具调用接口,将反应式动画与意图驱动行为结合,允许上游大语言模型注入语义控制。该框架可作为即插即用方案,将语音代理转化为交互式人形虚拟人。大量实验表明,本方法在动作质量和同步性上优于当前最优实时基线,同时具备现场部署所需灵活性。代码、预训练模型及视频见https://robinwitch.github.io/EchoAvatar-Page。

原文摘要 · Abstract (English)

Real-time synthesis of high-fidelity 3D character motion from audio is a pivotal component for next-generation interactive avatars and virtual assistants. However, most existing approaches are limited to offline processing of complete audio sequences or are constrained to specific domains, rarely handling both speech and music effectively. In this paper, we introduce a novel framework designed to generate continuous, coherent full-body motion from streaming speech and music with low latency. Central to our approach is a unified streaming architecture capable of synthesizing continuous motion from incremental audio inputs. We employ a robust training strategy that enforces strong audio dependency, allowing the model to seamlessly generalize across conversational speech and rhythmic music without requiring explicit domain labels or mode switching. Additionally, we explored Reinforcement Learning to refine the quality of online generation. Furthermore, we bridge reactive animation with intent-driven behavior via a tool-call interface that allows upstream Large Language Models to inject explicit semantic control. By combining this controllability with stream audio-driven synthesis, our framework serves as a plug-and-play solution for transforming voice agents into interactive humanoid avatars. Extensive experiments demonstrate that our method outperforms state-of-the-art realtime baselines in motion quality and synchronization while maintaining the flexibility required for live deployment. Our code, pre-trained models, and videos are available at https://robinwitch.github.io/EchoAvatar-Page.

虚拟人音频驱动实时生成

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。