用骨架辅助生成逼真同步的演讲手势视频,降低使用门槛。
Democratizing High-Fidelity Co-Speech Gesture Video Generation
- 以音频和参考图骨架为条件,融合特征预测动作
- 在405小时数据集上实现高质量音画同步生成
- 首个公开多类型高分辨率数据集,适合研究者复现
共语手势视频生成旨在合成与音频同步的说话人真实视频,包含面部表情与肢体动作。该任务因音频到视觉存在显著的一对多映射关系,且缺乏大规模公开数据集与高算力需求而困难。本文提出轻量级框架,利用2D全身骨骼作为高效辅助条件,将音频信号与视觉输出对齐。方法基于扩散模型,以细粒度音频段与从参考图像提取的骨架为输入,通过骨架-音频特征融合预测骨骼运动,确保严格音画同步与身体形态一致性。生成的骨骼序列结合说话人参考图像输入现有通用人体视频生成模型,合成高保真视频。为推动研究普及,我们构建了首个公开数据集CSG-405,含405小时高分辨率视频,覆盖71种语音类型,标注2D骨骼并涵盖多样化说话人背景。实验表明,本方法在视觉质量与同步性上超越现有技术,且具备跨说话人与场景泛化能力。代码、模型与数据集已公开于https://mpi-lab.github.io/Democratizing-CSG/
原文摘要 · Abstract (English)
Co-speech gesture video generation aims to synthesize realistic, audio-aligned videos of speakers, complete with synchronized facial expressions and body gestures. This task presents challenges due to the significant one-to-many mapping between audio and visual content, further complicated by the scarcity of large-scale public datasets and high computational demands. We propose a lightweight framework that utilizes 2D full-body skeletons as an efficient auxiliary condition to bridge audio signals with visual outputs. Our approach introduces a diffusion model conditioned on fine-grained audio segments and a skeleton extracted from the speaker's reference image, predicting skeletal motions through skeleton-audio feature fusion to ensure strict audio coordination and body shape consistency. The generated skeletons are then fed into an off-the-shelf human video generation model with the speaker's reference image to synthesize high-fidelity videos. To democratize research, we present CSG-405-the first public dataset with 405 hours of high-resolution videos across 71 speech types, annotated with 2D skeletons and diverse speaker demographics. Experiments show that our method exceeds state-of-the-art approaches in visual quality and synchronization while generalizing across speakers and contexts. Code, models, and CSG-405 are publicly released at https://mpi-lab.github.io/Democratizing-CSG/
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。