构建高质量人体视频数据集,提升生成效果。
OpenHumanVid: A Large-Scale High-Quality Dataset for Enhancing Human-Centric Video Generation
- 构建大规模人体视频数据集,含精准描述与运动条件。
- 使用该数据集训练后,人体视频生成质量显著提升。
- 适合研究人体生成、视频扩散模型的开发者使用。
视觉生成技术的进步大幅提升了视频数据集的规模和可用性,对训练高效视频生成模型至关重要。然而,高质量的人体中心视频数据集仍严重缺乏,制约了该领域发展。为此,我们提出 OpenHumanVid,一个大规模、高质量的人体中心视频数据集,具备精确详细的描述,涵盖人体外观与运动状态,并包含骨骼序列和语音音频等附加运动条件。为验证该数据集及训练策略的有效性,我们扩展了经典扩散变换器架构,并在该数据集上进行进一步预训练。结果表明:第一,使用大规模高质量数据集能显著提升生成人体视频的评估指标,同时保持通用视频生成任务的性能;第二,文本与人体外观、运动及面部动作的准确对齐是生成高质量视频的关键。基于此,仅通过简单扩展网络并在该数据集上训练,即可明显改善人体视频生成效果。
原文摘要 · Abstract (English)
Recent advancements in visual generation technologies have markedly increased the scale and availability of video datasets, which are crucial for training effective video generation models. However, a significant lack of high-quality, human-centric video datasets presents a challenge to progress in this field. To bridge this gap, we introduce OpenHumanVid, a large-scale and high-quality human-centric video dataset characterized by precise and detailed captions that encompass both human appearance and motion states, along with supplementary human motion conditions, including skeleton sequences and speech audio. To validate the efficacy of this dataset and the associated training strategies, we propose an extension of existing classical diffusion transformer architectures and conduct further pretraining of our models on the proposed dataset. Our findings yield two critical insights: First, the incorporation of a large-scale, high-quality dataset substantially enhances evaluation metrics for generated human videos while preserving performance in general video generation tasks. Second, the effective alignment of text with human appearance, human motion, and facial motion is essential for producing high-quality video outputs. Based on these insights and corresponding methodologies, the straightforward extended network trained on the proposed dataset demonstrates an obvious improvement in the generation of human-centric videos. Project page https://fudan-generative-vision.github.io/OpenHumanVid
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。