arXiv:2411.15355cs.CVcs.AI2024-11被引 13

统一高斯表示让自动驾驶场景重建兼容广角镜头。

UniGaussian: Driving Scene Reconstruction from Multiple Camera Models via Unified Gaussian Representations

  • 用仿射变换实现鱼眼相机的可微渲染,解决畸变问题。
  • 支持多源传感器融合,实时渲染且质量优于现有方法。
  • 适合需要多摄像头协同的自动驾驶仿真研究者。

城市场景重建对真实自动驾驶模拟器至关重要。现有方法虽实现逼真渲染,但多集中于针孔相机,忽略鱼眼相机。如何有效模拟鱼眼相机在驾驶场景中的表现仍是未解难题。本文提出UniGaussian,一种基于多相机模型的统一3D高斯表示方法,用于自动驾驶场景重建。首先,设计一种新的可微渲染方法,通过针对鱼眼相机模型的仿射变换扭曲3D高斯,解决3D高斯点阵与鱼眼相机不兼容的问题,该方法保持实时渲染并具备可微性。其次,构建新框架,通过仿射变换适配不同相机模型,并利用多模态监督(深度、语义、法线、LiDAR点云)正则化共享高斯,学习统一的3D高斯表示,实现对多传感器(针孔与鱼眼相机)和多模态数据的联合建模。实验表明,该方法在驾驶场景模拟中实现了更优的渲染质量与快速渲染速度。

原文摘要 · Abstract (English)

Urban scene reconstruction is crucial for real-world autonomous driving simulators. Although existing methods have achieved photorealistic reconstruction, they mostly focus on pinhole cameras and neglect fisheye cameras. In fact, how to effectively simulate fisheye cameras in driving scene remains an unsolved problem. In this work, we propose UniGaussian, a novel approach that learns a unified 3D Gaussian representation from multiple camera models for urban scene reconstruction in autonomous driving. Our contributions are two-fold. First, we propose a new differentiable rendering method that distorts 3D Gaussians using a series of affine transformations tailored to fisheye camera models. This addresses the compatibility issue of 3D Gaussian splatting with fisheye cameras, which is hindered by light ray distortion caused by lenses or mirrors. Besides, our method maintains real-time rendering while ensuring differentiability. Second, built on the differentiable rendering method, we design a new framework that learns a unified Gaussian representation from multiple camera models. By applying affine transformations to adapt different camera models and regularizing the shared Gaussians with supervision from different modalities, our framework learns a unified 3D Gaussian representation with input data from multiple sources and achieves holistic driving scene understanding. As a result, our approach models multiple sensors (pinhole and fisheye cameras) and modalities (depth, semantic, normal and LiDAR point clouds). Our experiments show that our method achieves superior rendering quality and fast rendering speed for driving scene simulation.

场景重建高斯表示鱼眼相机自动驾驶

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。