构建首个500万条人称视角视频数据集,推动虚拟现实视频生成
EgoVid-5M: A Large-Scale Video-Action Dataset for Egocentric Video Generation
- 构建500万条人称视角视频,含动作与运动控制标注
- 提出EgoDreamer模型,支持动作描述与运动信号联合生成
- 数据清洗确保画面连贯、动作一致,适合虚拟/增强现实研究
视频生成已成为模拟世界的重要工具,利用视觉数据复现真实环境。其中,以人类视角为中心的人称视角视频生成在虚拟现实、增强现实和游戏领域具有重要潜力。然而,由于人称视角动态变化、动作复杂多样、场景丰富多变,现有数据集难以有效应对这些挑战。为此,我们提出了EgoVid-5M,首个专为人类视角视频生成设计的高质量大规模数据集,包含500万条人称视角视频片段,并附有细粒度动作标注,包括运动学控制信号与高层文本描述。为保证数据质量,我们设计了复杂的数据清洗流程,确保帧间一致性、动作连贯性与运动平滑性。此外,我们引入EgoDreamer模型,可同时基于动作描述与运动控制信号生成人称视角视频。EgoVid-5M数据集、动作标注及所有清洗元数据将公开发布,以促进人称视角视频生成研究的发展。
原文摘要 · Abstract (English)
Video generation has emerged as a promising tool for world simulation, leveraging visual data to replicate real-world environments. Within this context, egocentric video generation, which centers on the human perspective, holds significant potential for enhancing applications in virtual reality, augmented reality, and gaming. However, the generation of egocentric videos presents substantial challenges due to the dynamic nature of egocentric viewpoints, the intricate diversity of actions, and the complex variety of scenes encountered. Existing datasets are inadequate for addressing these challenges effectively. To bridge this gap, we present EgoVid-5M, the first high-quality dataset specifically curated for egocentric video generation. EgoVid-5M encompasses 5 million egocentric video clips and is enriched with detailed action annotations, including fine-grained kinematic control and high-level textual descriptions. To ensure the integrity and usability of the dataset, we implement a sophisticated data cleaning pipeline designed to maintain frame consistency, action coherence, and motion smoothness under egocentric conditions. Furthermore, we introduce EgoDreamer, which is capable of generating egocentric videos driven simultaneously by action descriptions and kinematic control signals. The EgoVid-5M dataset, associated action annotations, and all data cleansing metadata will be released for the advancement of research in egocentric video generation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。