百万级视频数据集,融合3D几何与语义信息,支持感知与生成双重任务。
SceneScribe-1M: A Large-Scale Video Dataset with Comprehensive Geometric and Semantic Annotations

- 百万真实场景视频,含相机参数、深度图与3D点轨迹
- 在单目深度估计、动态追踪等任务上建立新基准
- 适合做3D视觉理解与可控视频生成的研究者使用
三维几何感知与视频生成的融合催生了对大规模、多模态视频数据的迫切需求。现有数据集虽在3D理解或视频生成方面有所进展,但在同时支持两个领域的大规模统一资源上仍存在显著空白。为此,我们推出SceneScribe-1M,一个包含一百万张野外视频的大型多模态数据集,每段视频均配有详细文本描述、精确相机参数、密集深度图及一致的3D点轨迹。我们通过在单目深度估计、场景重建、动态点跟踪以及带或不带相机控制的文本到视频生成等众多下游任务中建立基准,验证了SceneScribe-1M的通用性与价值。通过开源该数据集,我们旨在提供一个全面的评估基准,推动可同时感知动态三维世界并生成可控、逼真视频内容模型的发展。
原文摘要 · Abstract (English)
The convergence of 3D geometric perception and video synthesis has created an unprecedented demand for large-scale video data that is rich in both semantic and spatio-temporal information. While existing datasets have advanced either 3D understanding or video generation, a significant gap remains in providing a unified resource that supports both domains at scale. To bridge this chasm, we introduce SceneScribe-1M, a new large-scale, multi-modal video dataset. It comprises one million in-the-wild videos, each meticulously annotated with detailed textual descriptions, precise camera parameters, dense depth maps, and consistent 3D point tracks. We demonstrate the versatility and value of SceneScribe-1M by establishing benchmarks across a wide array of downstream tasks, including monocular depth estimation, scene reconstruction, and dynamic point tracking, as well as generative tasks such as text-to-video synthesis, with or without camera control. By open-sourcing SceneScribe-1M, we aim to provide a comprehensive benchmark and a catalyst for research, fostering the development of models that can both perceive the dynamic 3D world and generate controllable, realistic video content.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。