用统一模型实现真实视频的高效新视角生成
RayDer: Scalable Self-Supervised Novel View Synthesis from Real-World Video

- 单模型整合相机估计、场景重建与渲染,简化训练流程
- 在多规模数据和算力下呈清晰幂律增长,超越静态场景混合数据
- 零样本开集性能媲美顶尖监督方法,适合大规模视频应用
自监督新视角生成(NVS)虽有丰富视频数据,但因真实视频训练不稳定、多网络系统扩展性难预测而难以规模化。我们提出RayDer,一个统一的前馈式Transformer,将相机估计、场景重建与渲染整合为单一主干网络,使自监督NVS转化为可预测的单模型扩展问题。通过将时变内容建模为冗余因子,吸收动态变化,实现对无约束真实视频的稳定训练。重要的是,RayDer以静态场景为目标任务:动态内容仅作为可扩展的监督信号,不参与动态场景(4D)重建。在多个模型规模和数个数量级的数据量下,RayDer表现出清晰的幂律缩放特性,优于静态场景数据混合方案。在大量基准测试中,其零样本开集性能达到当前最先进监督方法水平。
原文摘要 · Abstract (English)
Self-supervised novel view synthesis (NVS) remains challenging to scale, despite the abundance of video data, largely due to the brittleness of training on realistic videos and the hard-to-predict scaling behavior of multi-network system designs. We introduce RayDer, a unified, feed-forward transformer that consolidates camera estimation, scene reconstruction, and rendering into a single backbone, turning self-supervised NVS into a well-posed single-model scaling problem. A minimal dynamic state, treated as a nuisance factor, absorbs time-varying content and enables stable training on unconstrained real-world video. Importantly, RayDer keeps static-scene NVS as its target task: dynamic content is leveraged purely as scalable supervision, not reconstructed as in dynamic-scene (4D) NVS. Across multiple model sizes and orders of magnitude in data, RayDer exhibits clean power-law scaling with data and compute, and outperforms static-scene data mixtures. On a large number of benchmarks, RayDer achieves strong zero-shot open-set performance competitive with state-of-the-art supervised approaches. Project Page: https://compvis.github.io/rayder
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。