arXiv:2608.09449cs.CV2026-08被引 1

Sekai2构建了可交互的世界建模数据集,支持长时间、多视角视频生成。

Sekai2: From World Exploration to Interactive World Modeling

论文配图:Sekai2: From World Exploration to Interactive World Modeling
图 1 · 摘自论文原文
  • 整合真实世界长视频与相机轨迹,实现多视角动态场景建模
  • 含64.9万段时序标注片段,51.4%视频达2分钟完整时长
  • 引入循环路径全景序列,助力长期空间记忆与几何一致性学习

视频世界模型需捕捉场景随时间与视角的变化。为支持长时程生成与相机控制,训练需依赖长视频、相机轨迹及时间对齐语义的联合数据。现有数据集通常缺乏三者之一:大规模网络视频虽视觉多样但无轨迹或时序文本;姿态标注数据集多为短程或重建导向。我们提出Sekai2,一个面向交互式世界建模的多源真实视频数据集,包含10,428个来源视频覆盖113个国家和地区,总计128,892个片段,总时长达2,826小时。在统一120秒划分下,43,594个片段达到完整时长,占总时长51.4%。每个片段均附带释放的相机轨迹和分层注释,分离主体运动、环境动态、静态场景与相机行为,共生成649,597个时序锚定段。关键创新在于引入982条非线性轨迹的全景序列,包含回访与重复观测,为学习持久场景表征、长期空间记忆与几何一致的世界模型提供关键监督。整体分析显示其具备完整的姿态与文本覆盖、广泛的地理与语义多样性、多样的相机轨迹以及高度非冗余的时间描述。这些特性使Sekai2成为长时程视频生成、相机可控合成与交互式世界模型预训练的可扩展资源。

原文摘要 · Abstract (English)

Video world models must capture how scenes evolve over time and across viewpoints. Training them for long-horizon generation and camera control therefore benefits from long videos paired with camera trajectories and temporally grounded semantics. Existing corpora rarely offer the three together: large-scale web video provides broad visual diversity but no trajectories or time-aligned text, while pose-annotated datasets are typically short-range or reconstruction-oriented. We introduce Sekai2, a multi-source real-world video dataset that carries the world-exploration footage of Sekai toward interactive world modeling. The release contains 128,892 clips totaling 2,826 hours from 10,428 source videos across 113 countries or regions, and is deliberately weighted toward sustained observation: under a common 120-second decomposition, 43,594 segments reach the full two minutes and account for 51.4% of all footage. Every clip includes a released camera trajectory and hierarchical annotations disentangling subject motion, environment dynamics, static scene content, and camera behavior, resulting in 649,597 temporally grounded segments. Crucially, we further introduce 982 panoramic sequences captured along non-linear trajectories with loops and revisits. These revisits provide repeated observations of the same locations across time and viewpoints, offering essential supervision for learning persistent scene representations, long-term spatial memory, and geometrically consistent world models. Corpus-scale analyses demonstrate complete pose-and-caption coverage, broad geographic and semantic diversity, varied camera trajectories, and highly non-redundant temporal descriptions. Together, these properties make Sekai2 a scalable resource for long-horizon video generation, camera-controllable synthesis, and interactive world-model pre-training.

世界模型视频生成长时程建模交互式建模

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。