用快速前向生成技术实现自动驾驶场景高保真4D重建。
ReconDrive: Fast Feed-Forward 4D Gaussian Splatting for Autonomous Driving Scene Reconstruction
- 改进3D基础模型,分离空间与外观预测以提升图像质量。
- 引入动静态时序融合策略,精准建模动态交通场景。
- 在nuScenes上速度远超优化方法,适合大规模仿真应用。
高保真视觉重建与新视角合成对自动驾驶闭环评估至关重要。尽管4D高斯点云(4DGS)在精度与效率间取得良好平衡,现有逐场景优化方法需耗时迭代,难以扩展至大规模城市环境。而当前前向方法常导致光照质量下降。为此,我们提出ReconDrive,一个基于VGGT的前向框架,可快速生成高保真4DGS。其核心改进包括:(1) 混合高斯预测头,解耦空间坐标与外观属性回归,克服通用基础特征的光照缺陷;(2) 静态-动态4D组合策略,通过速度建模显式捕捉时间运动,以表征复杂动态环境。在nuScenes数据集上,ReconDrive显著优于现有前向基线,在重建、新视角合成与3D感知任务中表现优异,性能接近逐场景优化方法,但速度提升数个数量级,为真实驾驶仿真提供可扩展且实用的解决方案。
原文摘要 · Abstract (English)
High-fidelity visual reconstruction and novel-view synthesis are essential for realistic closed-loop evaluation in autonomous driving. While 4D Gaussian Splatting (4DGS) offers a promising balance of accuracy and efficiency, existing per-scene optimization methods require costly iterative refinement, rendering them unscalable for extensive urban environments. Conversely, current feed-forward approaches often suffer from degraded photometric quality. To address these limitations, we propose ReconDrive, a feed-forward framework that leverages and extends the 3D foundation model VGGT for rapid, high-fidelity 4DGS generation. Our architecture introduces two core adaptations to tailor the foundation model to dynamic driving scenes: (1) Hybrid Gaussian Prediction Heads, which decouple the regression of spatial coordinates and appearance attributes to overcome the photometric deficiencies inherent in generalized foundation features; and (2) a Static-Dynamic 4D Composition strategy that explicitly captures temporal motion via velocity modeling to represent complex dynamic environments. Benchmarked on nuScenes, ReconDrive significantly outperforms existing feed-forward baselines in reconstruction, novel-view synthesis, and 3D perception. It achieves performance competitive with per-scene optimization while being orders of magnitude faster, providing a scalable and practical solution for realistic driving simulation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。