arXiv:2411.15482cs.CV2024-11CVPR被引 21

无需人工标注,通过自监督学习实现动态城市场景的4D重建与新视角生成。

SplatFlow: Self-Supervised Dynamic Gaussian Splatting in Neural Motion Flow Field for Autonomous Driving

  • 在神经运动流场中统一建模时空高斯表示,实现动态物体与静态背景分离。
  • 在Waymo和KITTI数据集上达到最新最优的图像重建与新视角合成效果。
  • 适合自动驾驶领域需要高效动态场景建模的研究者或工程师。

现有动态高斯点阵方法在复杂城市场景中依赖昂贵的人工标注对象级监督,限制了实际应用的可扩展性。本文提出SplatFlow,一种基于神经运动流场(NMFF)的自监督动态高斯点阵方法,可在无需追踪3D边界框的情况下学习4D时空表征,实现精准的动态场景重建与新视角RGB/深度/光流合成。SplatFlow设计统一框架,将时间相关的4D高斯表示融入NMFF,其中NMFF是一组隐式函数,用于建模激光雷达点与高斯点的时序运动,形成连续运动流场。借助NMFF,SplatFlow有效分解静态背景与动态物体,分别用3D和4D高斯基元表示;同时建模各4D高斯在时间上的对应关系,聚合时序特征以增强动态成分的跨视角一致性。此外,通过将2D基础模型特征蒸馏至4D时空表示,进一步提升动态物体识别能力。在Waymo与KITTI数据集上的全面评估表明,SplatFlow在动态城市场景的图像重建与新视角合成任务中均达到当前最优性能。

原文摘要 · Abstract (English)

Most existing Dynamic Gaussian Splatting methods for complex dynamic urban scenarios rely on accurate object-level supervision from expensive manual labeling, limiting their scalability in real-world applications. In this paper, we introduce SplatFlow, a Self-Supervised Dynamic Gaussian Splatting within Neural Motion Flow Fields (NMFF) to learn 4D space-time representations without requiring tracked 3D bounding boxes, enabling accurate dynamic scene reconstruction and novel view RGB/depth/flow synthesis. SplatFlow designs a unified framework to seamlessly integrate time-dependent 4D Gaussian representation within NMFF, where NMFF is a set of implicit functions to model temporal motions of both LiDAR points and Gaussians as continuous motion flow fields. Leveraging NMFF, SplatFlow effectively decomposes static background and dynamic objects, representing them with 3D and 4D Gaussian primitives, respectively. NMFF also models the correspondences of each 4D Gaussian across time, which aggregates temporal features to enhance cross-view consistency of dynamic components. SplatFlow further improves dynamic object identification by distilling features from 2D foundation models into 4D space-time representation. Comprehensive evaluations conducted on the Waymo and KITTI Datasets validate SplatFlow's state-of-the-art (SOTA) performance for both image reconstruction and novel view synthesis in dynamic urban scenarios.

动态建模自监督学习自动驾驶4D重建

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。