arXiv:2411.10061cs.GRcs.CV2024-11CVPR被引 81

用音频和姿态生成生动半身人像动画,减少控制条件

EchoMimicV2: Towards Striking, Simplified, and Semi-Body Human Animation

  • 通过音频-姿态动态调和策略,简化控制条件
  • 在半身数据稀缺下仍实现高细节与表情表现力
  • 适合需要轻量级、高质量人像动画的开发者

现有方法通常依赖音频、姿态或运动图作为控制条件,虽能生成逼真动画,但面临额外条件、复杂注入模块及仅限头部驱动等问题。为此,我们提出EchoMimicV2,一种基于新型音频-姿态动态调和策略(含姿态采样与音频扩散)的半身人像动画方法,有效提升面部与手势表现力,同时减少条件冗余。为弥补半身数据不足,引入头部分区注意力机制,可无缝融合头像数据训练且推理时无需,实现‘免费午餐’。此外,设计分阶段去噪损失,分别指导不同阶段的动作、细节与低层质量。还构建了首个半身人像动画评估基准。大量实验表明,该方法在定量与定性评价中均优于现有方法。

原文摘要 · Abstract (English)

Recent work on human animation usually involves audio, pose, or movement maps conditions, thereby achieves vivid animation quality. However, these methods often face practical challenges due to extra control conditions, cumbersome condition injection modules, or limitation to head region driving. Hence, we ask if it is possible to achieve striking half-body human animation while simplifying unnecessary conditions. To this end, we propose a half-body human animation method, dubbed EchoMimicV2, that leverages a novel Audio-Pose Dynamic Harmonization strategy, including Pose Sampling and Audio Diffusion, to enhance half-body details, facial and gestural expressiveness, and meanwhile reduce conditions redundancy. To compensate for the scarcity of half-body data, we utilize Head Partial Attention to seamlessly accommodate headshot data into our training framework, which can be omitted during inference, providing a free lunch for animation. Furthermore, we design the Phase-specific Denoising Loss to guide motion, detail, and low-level quality for animation in specific phases, respectively. Besides, we also present a novel benchmark for evaluating the effectiveness of half-body human animation. Extensive experiments and analyses demonstrate that EchoMimicV2 surpasses existing methods in both quantitative and qualitative evaluations.

人像动画音频驱动半身生成

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。