统一稀疏建模实现低延迟语音驱动的实时人脸与手势动画
UMo: Unified Sparse Motion Modeling for Real-Time Co-Speech Avatars

- 用统一框架处理文本、音频和动作标记,结合时空稀疏设计
- 在低延迟下仍保持精细语音-动作对齐,实现实时高质量动画
- 适合游戏、虚拟制作等需要高保真实时交互的场景
语音驱动的手势与面部动画是游戏、虚拟制作和互动媒体中生动数字人像的核心。现有方法或仅限单模态音频动作对齐,未能充分利用海量人体动作数据;或受多模态模型表示能力与吞吐量限制,难以兼顾高质量生成与实时性。我们提出UMo——一种面向实时语音同步虚拟人像的统一稀疏动作建模架构,将文本、音频与动作标记统一建模。通过空间稀疏的专家混合框架与时间稀疏的关键帧中心设计,高效实现稠密动作重建,生成具有时间连贯性与高保真度的面部表情与手势动画。此外,采用分阶段训练策略并引入针对性音频增强,提升声学多样性与语义一致性。实验表明,即使在严格延迟约束下,UMo仍能保持精细的语音-动作对齐,在低延迟与实时性能条件下显著优于现有方法,为高保真实时语音同步虚拟人像提供实用解决方案。
原文摘要 · Abstract (English)
Speech-driven gestures and facial animations are fundamental to expressive digital avatars in games, virtual production, and interactive media. However, existing methods are either limited to a single modality for audio motion alignment, failing to fully utilize the potential of massive human motion data, or are constrained by the representation ability and throughput of multimodal models, which makes it difficult to achieve high-quality motion generation or real-time performance. We present UMo, a unified sparse motion modeling architecture for real-time co-speech avatars, which processes text, audio, and motion tokens within a unified formulation. Leveraging a spatially sparse Mixture-of-Experts framework and a temporally sparse, keyframe-centric design, UMo efficiently performs real-time dense reconstruction, enabling temporally coherent and high-fidelity animation generation for both facial expressions and gestures. Furthermore, we implement a multi-stage training strategy with targeted audio augmentation to enhance acoustic diversity and semantic consistency. Consequently, UMo preserves fine-grained speech-motion alignment even under strict latency constraints. Extensive quantitative and qualitative evaluations show that UMo achieves better output quality under low latency and real-time performance constraints, offering a practical solution for high-fidelity real-time co-speech avatars.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。