通过多模态联合学习,实现自动驾驶场景的单视图高质量重建与零样本泛化。
ADGaussian: Generalizable Gaussian Splatting for Autonomous Driving via Multi-modal Joint Learning
- 融合视觉与稀疏激光雷达深度,联合优化图像与深度特征以预测高精度高斯点。
- 在Waymo和KITTI数据集上达到当前最优性能,新视角迁移零样本效果优异。
- 适合自动驾驶场景重建、多模态感知研究者使用。
我们提出一种名为ADGaussian的新方法,用于可泛化的街景三维重建。该方法仅需单视角输入即可实现高质量渲染。不同于以往主要关注几何优化的高斯点绘制方法,我们强调图像与深度特征联合优化对准确高斯点预测的重要性。为此,我们引入稀疏激光雷达深度作为额外模态,将高斯点预测建模为视觉信息与几何线索的联合学习框架。进一步地,提出多模态特征匹配策略与多尺度高斯解码模型,以增强多模态特征的联合优化,实现高效的多模态高斯学习。在Waymo和KITTI数据集上的大量实验表明,ADGaussian达到当前最优性能,并在新视角变换任务中展现出卓越的零样本泛化能力。
原文摘要 · Abstract (English)
We present a novel approach, termed ADGaussian, for generalizable street scene reconstruction. The proposed method enables high-quality rendering from merely single-view input. Unlike prior Gaussian Splatting methods that primarily focus on geometry refinement, we emphasize the importance of joint optimization of image and depth features for accurate Gaussian prediction. To this end, we first incorporate sparse LiDAR depth as an additional input modality, formulating the Gaussian prediction process as a joint learning framework of visual information and geometric clue. Furthermore, we propose a Multi-modal Feature Matching strategy coupled with a Multi-scale Gaussian Decoding model to enhance the joint refinement of multi-modal features, thereby enabling efficient multi-modal Gaussian learning. Extensive experiments on Waymo and KITTI demonstrate that our ADGaussian achieves state-of-the-art performance and exhibits superior zero-shot generalization capabilities in novel-view shifting.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。