基于射线空间的自回归建模,实现多视角驾驶场景的高保真视频生成。
RAYNOVA: Scale-Temporal Autoregressive World Modeling in Ray Space
- 采用尺度与时间双因果自回归框架,统一处理4D时空关系。
- 在nuScenes数据集上达到当前最优多视角视频生成效果。
- 无需显式3D场景表示,可泛化至新视角和相机配置。
世界基础模型旨在以物理合理的方式模拟真实世界的演变。不同于以往将空间与时间相关性分开处理的方法,我们提出RAYNOVA,一种几何无关的多视角驾驶场景世界模型,采用双因果自回归框架,遵循尺度与时间拓扑顺序,并利用全局注意力实现统一的4D时空推理。与依赖强3D几何先验的现有工作不同,RAYNOVA基于相对Plücker射线位置编码,构建跨视角、帧与尺度的各向同性时空表征,实现对多样化相机设置与自车运动的鲁棒泛化。我们进一步引入循环训练范式,缓解长时序视频生成中的分布漂移问题。RAYNOVA在nuScenes数据集上取得当前最优的多视角视频生成性能,同时具备更高吞吐量和强可控性,在多种输入条件下可泛化至新视角与相机配置,无需显式3D场景表示。代码将于https://raynova-ai.github.io/发布。
原文摘要 · Abstract (English)
World foundation models aim to simulate the evolution of the real world with physically plausible behavior. Unlike prior methods that handle spatial and temporal correlations separately, we propose RAYNOVA, a geometry-agonistic multiview world model for driving scenarios that employs a dual-causal autoregressive framework. It follows both scale-wise and temporal topological orders in the autoregressive process, and leverages global attention for unified 4D spatio-temporal reasoning. Different from existing works that impose strong 3D geometric priors, RAYNOVA constructs an isotropic spatio-temporal representation across views, frames, and scales based on relative Plücker-ray positional encoding, enabling robust generalization to diverse camera setups and ego motions. We further introduce a recurrent training paradigm to alleviate distribution drift in long-horizon video generation. RAYNOVA achieves state-of-the-art multi-view video generation results on nuScenes, while offering higher throughput and strong controllability under diverse input conditions, generalizing to novel views and camera configurations without explicit 3D scene representation. Our code will be released at https://raynova-ai.github.io/.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。