用音频直接生成表情与手势,让虚拟人说话更自然生动。
EMO2: End-Effector Guided Audio-Driven Avatar Video Generation
- 分两阶段生成:先从音频生成手部动作,再用扩散模型合成视频
- 在视觉质量和动作同步性上超越CyberHost和Vlogger等现有方法
- 适合需要高表达力的虚拟主播、教育动画等场景
本文提出一种新型音频驱动说话头生成方法,可同时生成高度表现力的面部表情与手部动作。不同于聚焦全身或半身姿态的现有方法,我们发现音频特征与全身动作之间的弱对应关系是关键限制。为此,将任务重新定义为两阶段流程:第一阶段直接从音频输入生成手部姿态,利用音频信号与手部运动间的强相关性;第二阶段采用扩散模型生成视频帧,并融入第一阶段生成的手部姿态,以合成逼真的面部表情与身体动作。实验结果表明,该方法在视觉质量与动作同步精度方面均优于当前最优方法,如CyberHost与Vlogger。本工作为音频驱动手势生成提供了新视角,并构建了一个生成富有表现力且自然的说话头动画的稳健框架。
原文摘要 · Abstract (English)
In this paper, we propose a novel audio-driven talking head method capable of simultaneously generating highly expressive facial expressions and hand gestures. Unlike existing methods that focus on generating full-body or half-body poses, we investigate the challenges of co-speech gesture generation and identify the weak correspondence between audio features and full-body gestures as a key limitation. To address this, we redefine the task as a two-stage process. In the first stage, we generate hand poses directly from audio input, leveraging the strong correlation between audio signals and hand movements. In the second stage, we employ a diffusion model to synthesize video frames, incorporating the hand poses generated in the first stage to produce realistic facial expressions and body movements. Our experimental results demonstrate that the proposed method outperforms state-of-the-art approaches, such as CyberHost and Vlogger, in terms of both visual quality and synchronization accuracy. This work provides a new perspective on audio-driven gesture generation and a robust framework for creating expressive and natural talking head animations.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。