无需标注,秒级重建大规模动态场景,支持未知物体泛化。
Flux4D: Flow-based Unsupervised 4D Reconstruction
- 直接预测3D高斯和运动动态,纯视觉无监督重建。
- 跨场景训练实现秒级重建,户外数据集上优于现有方法。
- 无需预训练模型或先验,适合大规模真实场景应用。
从视觉观测中重建大规模动态场景是计算机视觉中的核心挑战,对机器人与自动驾驶系统至关重要。尽管基于可微渲染的NeRF和3D高斯泼溅(3DGS)已实现逼真的三维重建,但其存在扩展性差且需标注以分离运动的问题。现有自监督方法虽尝试通过运动线索和几何先验消除显式标注,但仍受限于每场景优化及超参数敏感。本文提出Flux4D,一种简单且可扩展的大型动态场景4D重建框架。Flux4D通过仅使用光度损失并施加“尽可能静态”的正则化,直接从原始数据中预测3D高斯及其运动动态,实现完全无监督重建。该方法无需预训练模型或基础先验,仅通过多场景联合训练即可学习动态成分分解。实验表明,Flux4D在户外驾驶数据集上实现了显著更优的可扩展性、泛化能力与重建质量,可在数秒内完成动态场景重建,并有效拓展至大规模数据集与未知物体场景。
原文摘要 · Abstract (English)
Reconstructing large-scale dynamic scenes from visual observations is a fundamental challenge in computer vision, with critical implications for robotics and autonomous systems. While recent differentiable rendering methods such as Neural Radiance Fields (NeRF) and 3D Gaussian Splatting (3DGS) have achieved impressive photorealistic reconstruction, they suffer from scalability limitations and require annotations to decouple actor motion. Existing self-supervised methods attempt to eliminate explicit annotations by leveraging motion cues and geometric priors, yet they remain constrained by per-scene optimization and sensitivity to hyperparameter tuning. In this paper, we introduce Flux4D, a simple and scalable framework for 4D reconstruction of large-scale dynamic scenes. Flux4D directly predicts 3D Gaussians and their motion dynamics to reconstruct sensor observations in a fully unsupervised manner. By adopting only photometric losses and enforcing an "as static as possible" regularization, Flux4D learns to decompose dynamic elements directly from raw data without requiring pre-trained supervised models or foundational priors simply by training across many scenes. Our approach enables efficient reconstruction of dynamic scenes within seconds, scales effectively to large datasets, and generalizes well to unseen environments, including rare and unknown objects. Experiments on outdoor driving datasets show Flux4D significantly outperforms existing methods in scalability, generalization, and reconstruction quality.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。