同步生成3D动作与逼真视频,提升动作合理性与视觉质量
CoMoVi: Co-Generation of 3D Human Motions and Realistic Videos
- 通过双分支扩散模型实现动作与视频的联合生成
- 在自建数据集上实现高质量无外部参考的视频生成
- 适合动作生成、视频合成及跨模态研究者使用
本文发现3D人体动作与2D人体视频生成具有内在耦合性:3D动作提供合理性结构先验,预训练视频模型则赋予动作良好泛化能力。基于此,提出CoMoVi框架,在单一扩散去噪循环中同步生成3D动作与2D视频。由于3D动作与2D视频存在模态差异,我们设计了一种有效映射将3D动作转为对齐2D视频的表示,并构建双分支扩散模型,通过互特征交互与3D-2D交叉注意力实现动作与视频的协同生成。为训练与评估,我们构建了大规模真实世界人体视频数据集CoMoVi-Dataset,包含文本与动作标注,涵盖多样且具挑战性的动作。大量实验表明,该方法能生成高质量3D动作并具备更强泛化能力,同时可生成无需外部动作参考的高质量人体中心视频。
原文摘要 · Abstract (English)
In this paper, we find that the generation of 3D human motions and 2D human videos is intrinsically coupled. 3D motions provide the structural prior for plausibility and consistency in videos, while pre-trained video models offer strong generalization capabilities for motions. Based on this, we present CoMoVi, a co-generative framework that generates 3D human motions and videos synchronously within a single diffusion denoising loop. However, since the 3D human motions and the 2D human-centric videos have a modality gap between each other, we propose to project the 3D human motion into an effective 2D human motion representation that effectively aligns with the 2D videos. Then, we design a dual-branch diffusion model to couple human motion and the video generation process with mutual feature interaction and 3D-2D cross attentions. To train and evaluate our model, we curate CoMoVi-Dataset, a large-scale real-world human video dataset with text and motion annotations, covering diverse and challenging human motions. Extensive experiments demonstrate that our method generates high-quality 3D human motion with a better generalization ability and that our method can generate high-quality human-centric videos without external motion references.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。