让多人角色口型与情绪精准同步,生成高保真对话动画
HunyuanVideo-Avatar: High-Fidelity Audio-Driven Human Animation for Multiple Characters
- 用图像注入替代传统条件输入,保持角色一致性
- 通过情感参考图实现细粒度情绪控制,对齐音频与表情
- 多角色独立驱动,支持复杂场景下的真人级动画生成
近年来,语音驱动的人体动画取得显著进展,但在动态视频生成中保持角色一致性、实现角色与音频间精确情绪对齐,以及支持多角色语音驱动动画方面仍存在挑战。为此,我们提出 HunyuanVideo-Avatar,一种基于多模态扩散变换器(MM-DiT)的模型,可同时生成动态、情绪可控且支持多角色的对话视频。具体创新包括:(i) 设计角色图像注入模块,取代传统的加法式角色条件机制,消除训练与推理间的条件不匹配,保障动态动作与强角色一致性;(ii) 引入音频情感模块(AEM),从情感参考图中提取并迁移情感特征至生成视频,实现细粒度情绪风格控制;(iii) 提出面向人脸的音频适配器(FAA),通过潜在空间人脸掩码隔离音频驱动角色,利用交叉注意力实现多角色场景下的独立音频注入。上述创新使 HunyuanVideo-Avatar 在基准数据集和新提出的野外数据集上均超越现有方法,生成真实感强、动态沉浸的虚拟人物动画。
原文摘要 · Abstract (English)
Recent years have witnessed significant progress in audio-driven human animation. However, critical challenges remain in (i) generating highly dynamic videos while preserving character consistency, (ii) achieving precise emotion alignment between characters and audio, and (iii) enabling multi-character audio-driven animation. To address these challenges, we propose HunyuanVideo-Avatar, a multimodal diffusion transformer (MM-DiT)-based model capable of simultaneously generating dynamic, emotion-controllable, and multi-character dialogue videos. Concretely, HunyuanVideo-Avatar introduces three key innovations: (i) A character image injection module is designed to replace the conventional addition-based character conditioning scheme, eliminating the inherent condition mismatch between training and inference. This ensures the dynamic motion and strong character consistency; (ii) An Audio Emotion Module (AEM) is introduced to extract and transfer the emotional cues from an emotion reference image to the target generated video, enabling fine-grained and accurate emotion style control; (iii) A Face-Aware Audio Adapter (FAA) is proposed to isolate the audio-driven character with latent-level face mask, enabling independent audio injection via cross-attention for multi-character scenarios. These innovations empower HunyuanVideo-Avatar to surpass state-of-the-art methods on benchmark datasets and a newly proposed wild dataset, generating realistic avatars in dynamic, immersive scenarios.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。