arXiv:2506.18866cs.CVcs.AI2025-06被引 88

用音频生成全身动作视频,口型同步更准,还能精准控制造型。

OmniAvatar: Efficient Audio-Driven Avatar Video Generation with Adaptive Body Animation

  • 通过分层音频嵌入捕捉音效特征,提升口型与声音同步性。
  • 在面部和半身视频生成上优于现有模型,支持文本精确控制。
  • 适合播客、互动场景、歌唱等需要自然肢体动作的视频生成。

音频驱动的人体动画虽有进展,但多数方法仅关注面部动作,难以生成自然流畅的全身动画,且对精细生成的提示控制能力不足。为此,我们提出OmniAvatar——一种创新的音频驱动全身视频生成模型,显著提升唇形同步准确度与动作自然性。该模型采用像素级多层级音频嵌入策略,更好捕捉潜在空间中的音频特征,增强跨场景下的唇形同步。为在保持基础模型提示控制能力的同时有效融合音频特征,我们采用基于LoRA的训练方式。大量实验表明,OmniAvatar在面部及半身视频生成任务中均优于现有方法,支持通过文本实现精准控制,适用于播客、人际互动、动态场景及歌唱等多种应用。项目主页:https://omni-avatar.github.io/。

原文摘要 · Abstract (English)

Significant progress has been made in audio-driven human animation, while most existing methods focus mainly on facial movements, limiting their ability to create full-body animations with natural synchronization and fluidity. They also struggle with precise prompt control for fine-grained generation. To tackle these challenges, we introduce OmniAvatar, an innovative audio-driven full-body video generation model that enhances human animation with improved lip-sync accuracy and natural movements. OmniAvatar introduces a pixel-wise multi-hierarchical audio embedding strategy to better capture audio features in the latent space, enhancing lip-syncing across diverse scenes. To preserve the capability for prompt-driven control of foundation models while effectively incorporating audio features, we employ a LoRA-based training approach. Extensive experiments show that OmniAvatar surpasses existing models in both facial and semi-body video generation, offering precise text-based control for creating videos in various domains, such as podcasts, human interactions, dynamic scenes, and singing. Our project page is https://omni-avatar.github.io/.

音频驱动全身动画唇形同步文本控制

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。