arXiv:2409.01876cs.CVcs.AI2024-09被引 38

首个端到端音频驱动人体动画模型,支持零样本生成且保持手部完整

CyberHost: Taming Audio-driven Avatar Diffusion Model with Region Codebook Attention

  • 用区域码本注意力融合局部特征与运动先验,提升面部和手部动画质量
  • 引入身体运动图、手部清晰度评分等策略,显著改善动作自然性与细节
  • 适合需要高保真语音驱动角色动画的影视、虚拟人应用

基于扩散的视频生成技术已取得显著进展,推动了人体动画研究的爆发式增长。然而,现有方法大多局限于同模态驱动场景,跨模态人体动画仍鲜有探索。本文提出端到端音频驱动人体动画框架CyberHost,可实现手部完整性、身份一致性与自然动作表现。其核心是区域码本注意力机制,通过整合细粒度局部特征与学习到的运动模式先验,显著提升面部与手部动画生成质量。此外,我们设计了人体先验引导的训练策略,包括身体运动图、手部清晰度评分、姿态对齐参考特征及局部增强监督,进一步优化合成效果。据我们所知,CyberHost是首个支持零样本人体视频生成的端到端音频驱动扩散模型。大量实验表明,该模型在定量与定性指标上均优于已有方法。

原文摘要 · Abstract (English)

Diffusion-based video generation technology has advanced significantly, catalyzing a proliferation of research in human animation. However, the majority of these studies are confined to same-modality driving settings, with cross-modality human body animation remaining relatively underexplored. In this paper, we introduce, an end-to-end audio-driven human animation framework that ensures hand integrity, identity consistency, and natural motion. The key design of CyberHost is the Region Codebook Attention mechanism, which improves the generation quality of facial and hand animations by integrating fine-grained local features with learned motion pattern priors. Furthermore, we have developed a suite of human-prior-guided training strategies, including body movement map, hand clarity score, pose-aligned reference feature, and local enhancement supervision, to improve synthesis results. To our knowledge, CyberHost is the first end-to-end audio-driven human diffusion model capable of facilitating zero-shot video generation within the scope of human body. Extensive experiments demonstrate that CyberHost surpasses previous works in both quantitative and qualitative aspects.

音频驱动扩散模型虚拟人手部动画

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。