构建可扩展的三维世界探索数据集,支持多视角与丰富标注。
WorldRover: A Scalable Synthetic Video Data Engine for World Exploration with Rich Annotations

- 基于Unreal引擎生成带完整轨迹与几何的长时序视频。
- 提供含深度、相机位姿、动作信号等1000万帧标注数据。
- 适合训练交互式场景理解与动态建模的AI模型。
生成可探索世界的模型需要视频配以超越RGB的信息:相机运动、场景几何、时间对应关系,以及交互模型所需的控制信号。真实采集可提供部分信号,但密集几何与长程对应通常依赖估计或专用设备。渲染可直接提供这些信息,但现有合成资源极少在同一帧中同时包含上述内容,且难以支持视角与外观的可控变化。本文提出WorldRover,一个用于生成富含标注的艺术家构建环境长程探索数据的引擎。核心为Unreal Engine管道,可执行并离线渲染分钟级路径,完整保留轨迹与场景几何。同一探索可从第一人称、第三人称及360全景视角重播,且在不同环境状态或中性白色材质下呈现。利用该引擎,我们构建了WorldRover-10M,其序列在整段探索中配对了RGB、度量深度、相机轨迹和轨迹推导出的动作信号。第三人称子集额外提供密集光流、长程2D/3D点轨迹(含可见性)以及独立于相机轨迹的角色轨迹。该引擎支持从第一人称、第三人称及360全景视角渲染,可在不同环境状态或中性白材质下进行,同时保持路线与场景几何一致。WorldRover将长时序世界探索转化为可扩展的数据生成问题,为需构建、维持并反复访问一致世界表征的模型提供监督信号。
原文摘要 · Abstract (English)
Learning to generate or reconstruct explorable worlds requires video paired with more than RGB: camera motion, scene geometry, temporal correspondence and, for interactive models, control signals. Real capture can provide some of these signals, but dense geometry and long-range correspondence usually rely on estimation or specialised instrumentation. Rendering provides these quantities directly, yet existing synthetic resources rarely combine them on the same frames while also supporting controlled changes of viewpoint and appearance. We introduce WorldRover, a data engine for generating richly annotated, long-range explorations of artist-built environments. At its core, WorldRover-Engine is an Unreal Engine pipeline that executes and offline-renders minute-scale routes while preserving their full trajectories and scene geometry. The same exploration can be replayed from first-person, third-person, and 360 panoramic cameras under different environmental states. Using WorldRover-Engine, we construct WorldRover-10M, whose sequences pair RGB with metric depth, camera trajectories, and trajectory-derived action signals throughout each exploration. Third-person subsets additionally provide dense optical flow, long-range 2D/3D point tracks with visibility, and a character trajectory distinct from the camera trajectory. The engine can render a traversal from first-person, third-person and 360 panoramic viewpoints, under different environmental states or with a neutral white material, while preserving the route and scene geometry. WorldRover therefore turns long-horizon world exploration into a scalable data-generation problem, providing supervision for models that must build, maintain, and revisit coherent representations of an explorable world.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。