arXiv:2506.18890cs.CV2025-06NeurIPS被引 11

4D-LRM可任意时间任意视角重建动态物体,一次前向计算仅需1.5秒。

4D-LRM: Large Space-Time Reconstruction Model From and To Any View at Any Time

  • 统一学习时空表示,直接从多视角图像预测4D高斯原语。
  • 单次前向传播重建24帧序列,耗时不足1.5秒(A100 GPU)。
  • 适用于新物体、跨时间插值及多样相机设置,适合实时4D重建场景。

能否将4D预训练扩展至学习通用时空表征,实现从若干时刻的任意视角重建任意视角与任意时间?我们通过4D-LRM给出了肯定回答:这是首个大规模4D重建模型,能处理非受限视角和时间戳输入,并渲染任意新视图-时间组合。与以往基于优化、几何或生成的方法相比,4D-LRM克服了效率、泛化性与保真度的局限,学习统一时空表示,直接从跨时间的带姿态图像令牌预测每像素4D高斯原语,实现理论上无限帧率的快速高质量渲染。实验表明,时空预训练规模提升可实现精确高效的4D重建。4D-LRM展现出对新物体的泛化能力、跨时间插值能力以及对多样化相机设置的适应性。在单张A100 GPU上,一次前向传播即可完成24帧序列重建,耗时低于1.5秒。

原文摘要 · Abstract (English)

Can we scale 4D pretraining to learn general space-time representations that reconstruct an object from a few views at some times to any view at any time? We provide an affirmative answer with 4D-LRM, the first large-scale 4D reconstruction model that takes input from unconstrained views and timestamps and renders arbitrary novel view-time combinations. Unlike prior 4D approaches, e.g., optimization-based, geometry-based, or generative, that struggle with efficiency, generalization, or faithfulness, 4D-LRM learns a unified space-time representation and directly predicts per-pixel 4D Gaussian primitives from posed image tokens across time, enabling fast, high-quality rendering at, in principle, infinite frame rate. Our results demonstrate that scaling spatiotemporal pretraining enables accurate and efficient 4D reconstruction. We show that 4D-LRM generalizes to novel objects, interpolates across time, and handles diverse camera setups. It reconstructs 24-frame sequences in one forward pass with less than 1.5 seconds on a single A100 GPU.

4D重建时空建模高斯渲染实时渲染

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。