用音频和运动掩码生成更自然的同步手势视频
MMGT: Motion Mask Guided Two-Stage Network for Co-Speech Gesture Video Generation
- 分两阶段生成:先从音频生成姿态和运动掩码,再融合掩码控制细节
- 在Lip-sync、手部动作和整体视频质量上优于现有方法
- 适合需要高精度手势生成的应用,如虚拟主播或数字人
共言语手势视频生成旨在从音频驱动的静态图像生成生动的说话视频,但因身体部位运动幅度、音频相关性和细节特征差异大而极具挑战。仅依赖音频常难以捕捉大幅手势,导致明显伪影和失真。现有方法多添加额外先验输入,限制了实用性。本文提出运动掩码引导的两阶段网络(MMGT),结合音频、运动掩码及由音频生成的姿态视频,协同生成同步的语音手势视频。第一阶段通过空间掩码引导的Audio2Pose网络(SMGA)从音频生成高质量姿态视频与运动掩码,有效捕捉面部和手势等关键区域的大范围运动。第二阶段将运动掩码融合进稳定扩散视频生成模型中的分层音频注意力机制(MM-HAA),解决传统方法在细粒度运动生成与区域细节控制上的不足,确保上半身视频具有高质量纹理和精确运动。实验表明,该方法在视频质量、口型同步和手部动作方面均有显著提升。模型与代码已开源于https://github.com/SIA-IDE/MMGT。
原文摘要 · Abstract (English)
Co-Speech Gesture Video Generation aims to generate vivid speech videos from audio-driven still images, which is challenging due to the diversity of body parts in terms of motion amplitude, audio relevance, and detailed features. Relying solely on audio as the control signal often fails to capture large gesture movements in videos, resulting in more noticeable artifacts and distortions. Existing approaches typically address this issue by adding extra prior inputs, but this can limit the practical application of the task. Specifically, we propose a Motion Mask-Guided Two-Stage Network (MMGT) that uses audio, along with motion masks and pose videos generated from the audio signal, to jointly generate synchronized speech gesture videos. In the first stage, the Spatial Mask-Guided Audio2Pose Generation (SMGA) Network generates high-quality pose videos and motion masks from audio, effectively capturing large movements in key regions such as the face and gestures. In the second stage, we integrate Motion Masked Hierarchical Audio Attention (MM-HAA) into the Stabilized Diffusion Video Generation model, addressing limitations in fine-grained motion generation and region-specific detail control found in traditional methods. This ensures high-quality, detailed upper-body videos with accurate textures and motion. Evaluations demonstrate improvements in video quality, lip-sync, and hand gestures. The model and code are available at https://github.com/SIA-IDE/MMGT.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。