构建首个全球范围第一视角视频数据集,支持世界探索型视频生成训练
Sekai: A Video Dataset towards World Exploration
- 收集超5000小时跨100国750城的步行/无人机视角视频
- 涵盖位置、天气、人流密度等多维度标注,支持真实世界建模
- 专为交互式世界探索设计,适合视频生成与具身智能研究者
视频生成技术取得显著进展,有望成为交互式世界探索的基础。然而,现有视频生成数据集在世界探索训练中存在局限:地点有限、时长较短、场景静态,且缺乏探索行为与世界结构的标注。本文提出Sekai(日语意为“世界”),一个高质量的第一人称视角全球视频数据集,包含超过5000小时来自100多个国家和地区、750个城市的步行或无人机视角(FPV和UVA)视频。我们开发了高效工具链,用于视频采集、预处理及标注,涵盖位置、场景、天气、人群密度、字幕和相机轨迹等信息。全面分析与实验表明,该数据集具有规模大、多样性高、标注质量优,并能有效训练视频生成模型。我们认为Sekai将推动视频生成与世界探索领域的发展,并激发重要应用。项目页面:https://lixsp11.github.io/sekai-project/
原文摘要 · Abstract (English)
Video generation techniques have made remarkable progress, promising to be the foundation of interactive world exploration. However, existing video generation datasets are not well-suited for world exploration training as they suffer from some limitations: limited locations, short duration, static scenes, and a lack of annotations about exploration and the world. In this paper, we introduce Sekai (meaning "world" in Japanese), a high-quality first-person view worldwide video dataset with rich annotations for world exploration. It consists of over 5,000 hours of walking or drone view (FPV and UVA) videos from over 100 countries and regions across 750 cities. We develop an efficient and effective toolbox to collect, pre-process and annotate videos with location, scene, weather, crowd density, captions, and camera trajectories. Comprehensive analyses and experiments demonstrate the dataset's scale, diversity, annotation quality, and effectiveness for training video generation models. We believe Sekai will benefit the area of video generation and world exploration, and motivate valuable applications. The project page is https://lixsp11.github.io/sekai-project/.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。