构建大规模带空间标注的视频数据集,助力3D视觉模型泛化能力提升。
SpatialVID: A Large-Scale Video Dataset with Spatial Annotations
- 采集21000小时原始视频,经筛选生成7089小时动态内容
- 每帧包含相机位姿、深度图、运动掩码等密集空间标注
- 适合研究视频理解、三维重建与动态场景建模的学者使用
空间智能领域虽取得显著进展,但现有模型的可扩展性与真实世界适应性仍受限于高质量训练数据的稀缺。尽管已有数据集提供相机位姿信息,但普遍存在规模小、多样性不足、标注不丰富等问题,尤其缺乏真实动态场景下的真实相机运动标注。为此,我们构建了SpatialVID,一个包含多样化场景、复杂相机运动的大规模野外视频数据集,具备密集的3D标注。我们采集超过21,000小时原始视频,通过分层过滤流程处理为270万段视频剪辑,总计7,089小时动态内容。后续标注流程为这些片段增加了详细的时空与语义信息,包括每帧相机位姿、深度图、动态掩码、结构化描述及序列化运动指令。对SpatialVID的数据统计分析表明,其丰富的多样性和标注密度能直接促进模型泛化性能提升,是视频与三维视觉研究社区的重要资产。
原文摘要 · Abstract (English)
Significant progress has been made in spatial intelligence, spanning both spatial reconstruction and world exploration. However, the scalability and real-world fidelity of current models remain severely constrained by the scarcity of large-scale, high-quality training data. While several datasets provide camera pose information, they are typically limited in scale, diversity, and annotation richness, particularly for real-world dynamic scenes with ground-truth camera motion. To this end, we collect SpatialVID, a dataset consists of a large corpus of in-the-wild videos with diverse scenes, camera movements and dense 3D annotations such as per-frame camera poses, depth, and motion instructions. Specifically, we collect more than 21,000 hours of raw videos, and process them into 2.7 million clips through a hierarchical filtering pipeline, totaling 7,089 hours of dynamic content. A subsequent annotation pipeline enriches these clips with detailed spatial and semantic information, including camera poses, depth maps, dynamic masks, structured captions, and serialized motion instructions. Analysis of SpatialVID's data statistics reveals a richness and diversity that directly fosters improved model generalization and performance, establishing it as a key asset for the video and 3D vision research community.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。