arXiv:2501.14646cs.CV2025-01IJCAI被引 7

实时生成高保真语音驱动的虚拟人,面部与身体动作同步自然。

SyncAnimation: A Real-Time End-to-End Framework for Audio-Driven Human Pose and Talking Head Animation

  • 基于NeRF框架,融合语音到姿态与表情的双同步机制。
  • 实现音频同步的头部、上身与口型动态,全程低延迟。
  • 适合直播、虚拟助手等对实时性与视觉质量要求高的场景。

语音驱动的虚拟人生成仍面临重大挑战。现有方法通常计算开销大,面部细节不足,难以满足高实时性与高质量视觉表现的应用需求。尽管部分方法可实现口型同步,但在无声时段仍存在面部表情与上半身动作不一致的问题。本文提出SyncAnimation,首个基于NeRF的端到端实时语音驱动虚拟人生成框架,通过结合通用语音到姿态匹配与语音到表情同步机制,实现高精度的姿态与表情生成,逐步构建出与音频同步的上半身、头部及口型变化。此外,高同步人像渲染器确保了头身融合自然,并实现精准的口型同步。项目主页见:https://syncanimation.github.io

原文摘要 · Abstract (English)

Generating talking avatar driven by audio remains a significant challenge. Existing methods typically require high computational costs and often lack sufficient facial detail and realism, making them unsuitable for applications that demand high real-time performance and visual quality. Additionally, while some methods can synchronize lip movement, they still face issues with consistency between facial expressions and upper body movement, particularly during silent periods. In this paper, we introduce SyncAnimation, the first NeRF-based method that achieves audio-driven, stable, and real-time generation of speaking avatar by combining generalized audio-to-pose matching and audio-to-expression synchronization. By integrating AudioPose Syncer and AudioEmotion Syncer, SyncAnimation achieves high-precision poses and expression generation, progressively producing audio-synchronized upper body, head, and lip shapes. Furthermore, the High-Synchronization Human Renderer ensures seamless integration of the head and upper body, and achieves audio-sync lip. The project page can be found at https://syncanimation.github.io

虚拟人语音驱动实时生成NeRF

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。